Prompt Chain

Integrate OpenAI MCP with Promptfoo Responses

A promptfoo example for evaluating OpenAI's MCP tool integration, including auth and approval workflows.

Works with openai

92
Spark score
out of 100
Updated 27 days ago
Version 0.121.18

Add to Favorites

Why it matters

Leverage OpenAI's Model Context Protocol (MCP) within promptfoo for enhanced response handling and integration. This asset streamlines the process of connecting and utilizing advanced AI model features.

Outcomes

What it gets done

01

Demonstrate OpenAI MCP integration with promptfoo.

02

Utilize promptfoo's Responses API for model interactions.

03

Facilitate context management for AI model responses.

Install

Add it to your toolbox

Run in your project directory:

curl -fsSL https://spark.entire.vc/get/pfoo-openai-mcp | bash

Steps

Steps in the chain

01
Set up environment variables
02
Run individual examples
03
View results

Overview

Openai Mcp

A promptfoo example for evaluating OpenAI's MCP tool integration through the Responses API, covering public, authenticated, and approval-workflow MCP server scenarios. Use it to evaluate MCP tool call correctness, authentication, and approval behavior. The authenticated scenario needs a STRIPE_API_KEY in addition to an OpenAI key.

What it does

This promptfoo example demonstrates OpenAI's Model Context Protocol (MCP) integration with the Responses API - an open protocol that standardizes how applications give tools and context to LLMs, letting a model call remote MCP servers to search repositories, access APIs, and more. It covers three scenarios: basic MCP usage against the public DeepWiki server for searching GitHub repositories, authenticated MCP against Stripe's MCP server, and configurable approval workflows for MCP tool calls.

When to use - and when NOT to

Use this example when you want to evaluate how an OpenAI model uses remote MCP tools, whether public and unauthenticated or requiring credentials, and how different approval settings affect tool execution. The authenticated example specifically needs a STRIPE_API_KEY in addition to OPENAI_API_KEY; the basic and approval examples only need the OpenAI key.

Inputs and outputs

MCP tools are configured with a server_label, server_url, and a require_approval setting:

tools:
  - type: mcp
    server_label: deepwiki
    server_url: https://mcp.deepwiki.com/mcp
    require_approval: never

Authenticated servers add a headers block (for example, an Authorization Bearer token interpolated from STRIPE_API_KEY); allowed_tools restricts which tools on a server are callable; and require_approval can be scoped per tool name instead of applying to every tool uniformly, for selective approval rather than an all-or-nothing setting.

Integrations

Assertions validate MCP interactions at several levels: is-valid-openai-tools-call checks both function and MCP tool calls are well-formed, contains/not-contains checks confirm MCP tool results appeared without MCP errors, and a weighted combination of tool-call validity, content matching, and an llm-rubric judgment can score response quality across multiple dimensions at once. Weighted assertions let you combine mechanical checks (did the tool call validate, does the response contain expected content) with a softer llm-rubric judgment of overall response quality, rather than relying on any single check alone. Run the three example configs individually - promptfooconfig.yaml for basic usage, promptfooconfig.authenticated.yaml for the Stripe example, and promptfooconfig.approval.yaml for approval-workflow variants - with npx promptfoo eval -c <config>.yaml, then npx promptfoo view to inspect results.

Who it's for

Teams building OpenAI-based agents that call remote MCP servers who need to validate tool call correctness, authentication handling, and approval-workflow behavior before relying on MCP tools in production. Further reference material is linked from the example itself: OpenAI's own MCP documentation, the Model Context Protocol specification, promptfoo's OpenAI provider docs, and the public MCP Server Registry for discovering additional servers to test against beyond DeepWiki and Stripe.

Source README

openai-mcp (OpenAI MCP Integration)

This example demonstrates how to use OpenAI's Model Context Protocol (MCP) integration with the Responses API in promptfoo.

You can run this example with:

npx promptfoo@latest init --example openai-mcp
cd openai-mcp

What is MCP?

Model Context Protocol (MCP) is an open protocol that standardizes how applications provide tools and context to LLMs. OpenAI's MCP integration allows models to use remote MCP servers to perform tasks like searching repositories, accessing APIs, and more.

Environment Variables

This example requires the following environment variables:

  • OPENAI_API_KEY - Your OpenAI API key from the OpenAI platform
  • STRIPE_API_KEY - Your Stripe API key (for authenticated examples only)

You can set these in a .env file or directly in your environment.

Running the Examples

  1. Set up environment variables:

    export OPENAI_API_KEY="your_openai_api_key"
    export STRIPE_API_KEY="your_stripe_api_key"  # For authenticated example only
    
  2. Run individual examples:

    # Basic MCP integration
    npx promptfoo eval -c promptfooconfig.yaml
    
    # Authenticated MCP servers
    npx promptfoo eval -c promptfooconfig.authenticated.yaml
    
    # Approval workflow examples
    npx promptfoo eval -c promptfooconfig.approval.yaml
    
  3. View results:

    npx promptfoo view
    

Examples

Basic MCP Usage (promptfooconfig.yaml)

Demonstrates basic MCP integration using the public DeepWiki MCP server to search GitHub repositories.

Authenticated MCP (promptfooconfig.authenticated.yaml)

Shows how to use MCP servers that require authentication, using Stripe as an example.

Approval Workflows (promptfooconfig.approval.yaml)

Demonstrates different approval settings for MCP tool usage.

MCP Configuration

Basic Configuration

tools:
  - type: mcp
    server_label: deepwiki
    server_url: https://mcp.deepwiki.com/mcp
    require_approval: never

With Authentication

tools:
  - type: mcp
    server_label: stripe
    server_url: https://mcp.stripe.com
    headers:
      Authorization: 'Bearer ${STRIPE_API_KEY}'
    require_approval: never

Tool Filtering

tools:
  - type: mcp
    server_label: deepwiki
    server_url: https://mcp.deepwiki.com/mcp
    allowed_tools: ['ask_question', 'read_wiki_structure']
    require_approval: never

Selective Approval

tools:
  - type: mcp
    server_label: deepwiki
    server_url: https://mcp.deepwiki.com/mcp
    require_approval:
      never:
        tool_names: ['ask_question']

Assertion Patterns

The examples demonstrate assertion patterns for validating MCP tool interactions:

Enhanced OpenAI Tools Validation

assert:
  - type: is-valid-openai-tools-call # Validates both function and MCP tools

MCP-Specific Content Validation

assert:
  - type: contains
    value: 'MCP Tool Result' # Verify MCP tools were used
  - type: not-contains
    value: 'MCP Tool Error' # Ensure no MCP errors occurred

Multi-layered Validation

assert:
  - type: is-valid-openai-tools-call
    weight: 0.4
  - type: contains-any
    value: ['expected', 'content']
    weight: 0.3
  - type: llm-rubric
    value: 'Response quality criteria'
    weight: 0.3

Learn More

FAQ

Common questions

Discussion

Questions & comments · 0

Sign In Sign in to leave a comment.