Integrate OpenAI MCP with Promptfoo Responses
A promptfoo example for evaluating OpenAI's MCP tool integration, including auth and approval workflows.
Why it matters
Leverage OpenAI's Model Context Protocol (MCP) within promptfoo for enhanced response handling and integration. This asset streamlines the process of connecting and utilizing advanced AI model features.
Outcomes
What it gets done
Demonstrate OpenAI MCP integration with promptfoo.
Utilize promptfoo's Responses API for model interactions.
Facilitate context management for AI model responses.
Install
Add it to your toolbox
Run in your project directory:
curl -fsSL https://spark.entire.vc/get/pfoo-openai-mcp | bash Steps
Steps in the chain
Overview
Openai Mcp
A promptfoo example for evaluating OpenAI's MCP tool integration through the Responses API, covering public, authenticated, and approval-workflow MCP server scenarios. Use it to evaluate MCP tool call correctness, authentication, and approval behavior. The authenticated scenario needs a STRIPE_API_KEY in addition to an OpenAI key.
What it does
This promptfoo example demonstrates OpenAI's Model Context Protocol (MCP) integration with the Responses API - an open protocol that standardizes how applications give tools and context to LLMs, letting a model call remote MCP servers to search repositories, access APIs, and more. It covers three scenarios: basic MCP usage against the public DeepWiki server for searching GitHub repositories, authenticated MCP against Stripe's MCP server, and configurable approval workflows for MCP tool calls.
When to use - and when NOT to
Use this example when you want to evaluate how an OpenAI model uses remote MCP tools, whether public and unauthenticated or requiring credentials, and how different approval settings affect tool execution. The authenticated example specifically needs a STRIPE_API_KEY in addition to OPENAI_API_KEY; the basic and approval examples only need the OpenAI key.
Inputs and outputs
MCP tools are configured with a server_label, server_url, and a require_approval setting:
tools:
- type: mcp
server_label: deepwiki
server_url: https://mcp.deepwiki.com/mcp
require_approval: never
Authenticated servers add a headers block (for example, an Authorization Bearer token interpolated from STRIPE_API_KEY); allowed_tools restricts which tools on a server are callable; and require_approval can be scoped per tool name instead of applying to every tool uniformly, for selective approval rather than an all-or-nothing setting.
Integrations
Assertions validate MCP interactions at several levels: is-valid-openai-tools-call checks both function and MCP tool calls are well-formed, contains/not-contains checks confirm MCP tool results appeared without MCP errors, and a weighted combination of tool-call validity, content matching, and an llm-rubric judgment can score response quality across multiple dimensions at once. Weighted assertions let you combine mechanical checks (did the tool call validate, does the response contain expected content) with a softer llm-rubric judgment of overall response quality, rather than relying on any single check alone. Run the three example configs individually - promptfooconfig.yaml for basic usage, promptfooconfig.authenticated.yaml for the Stripe example, and promptfooconfig.approval.yaml for approval-workflow variants - with npx promptfoo eval -c <config>.yaml, then npx promptfoo view to inspect results.
Who it's for
Teams building OpenAI-based agents that call remote MCP servers who need to validate tool call correctness, authentication handling, and approval-workflow behavior before relying on MCP tools in production. Further reference material is linked from the example itself: OpenAI's own MCP documentation, the Model Context Protocol specification, promptfoo's OpenAI provider docs, and the public MCP Server Registry for discovering additional servers to test against beyond DeepWiki and Stripe.
Source README
openai-mcp (OpenAI MCP Integration)
This example demonstrates how to use OpenAI's Model Context Protocol (MCP) integration with the Responses API in promptfoo.
You can run this example with:
npx promptfoo@latest init --example openai-mcp
cd openai-mcp
What is MCP?
Model Context Protocol (MCP) is an open protocol that standardizes how applications provide tools and context to LLMs. OpenAI's MCP integration allows models to use remote MCP servers to perform tasks like searching repositories, accessing APIs, and more.
Environment Variables
This example requires the following environment variables:
OPENAI_API_KEY- Your OpenAI API key from the OpenAI platformSTRIPE_API_KEY- Your Stripe API key (for authenticated examples only)
You can set these in a .env file or directly in your environment.
Running the Examples
Set up environment variables:
export OPENAI_API_KEY="your_openai_api_key" export STRIPE_API_KEY="your_stripe_api_key" # For authenticated example onlyRun individual examples:
# Basic MCP integration npx promptfoo eval -c promptfooconfig.yaml # Authenticated MCP servers npx promptfoo eval -c promptfooconfig.authenticated.yaml # Approval workflow examples npx promptfoo eval -c promptfooconfig.approval.yamlView results:
npx promptfoo view
Examples
Basic MCP Usage (promptfooconfig.yaml)
Demonstrates basic MCP integration using the public DeepWiki MCP server to search GitHub repositories.
Authenticated MCP (promptfooconfig.authenticated.yaml)
Shows how to use MCP servers that require authentication, using Stripe as an example.
Approval Workflows (promptfooconfig.approval.yaml)
Demonstrates different approval settings for MCP tool usage.
MCP Configuration
Basic Configuration
tools:
- type: mcp
server_label: deepwiki
server_url: https://mcp.deepwiki.com/mcp
require_approval: never
With Authentication
tools:
- type: mcp
server_label: stripe
server_url: https://mcp.stripe.com
headers:
Authorization: 'Bearer ${STRIPE_API_KEY}'
require_approval: never
Tool Filtering
tools:
- type: mcp
server_label: deepwiki
server_url: https://mcp.deepwiki.com/mcp
allowed_tools: ['ask_question', 'read_wiki_structure']
require_approval: never
Selective Approval
tools:
- type: mcp
server_label: deepwiki
server_url: https://mcp.deepwiki.com/mcp
require_approval:
never:
tool_names: ['ask_question']
Assertion Patterns
The examples demonstrate assertion patterns for validating MCP tool interactions:
Enhanced OpenAI Tools Validation
assert:
- type: is-valid-openai-tools-call # Validates both function and MCP tools
MCP-Specific Content Validation
assert:
- type: contains
value: 'MCP Tool Result' # Verify MCP tools were used
- type: not-contains
value: 'MCP Tool Error' # Ensure no MCP errors occurred
Multi-layered Validation
assert:
- type: is-valid-openai-tools-call
weight: 0.4
- type: contains-any
value: ['expected', 'content']
weight: 0.3
- type: llm-rubric
value: 'Response quality criteria'
weight: 0.3
Learn More
FAQ
Common questions
Discussion
Questions & comments · 0
Sign In Sign in to leave a comment.