3,603 tools found
A promptfoo example for evaluating multi-agent Python openai-agents SDK workflows with trace-level assertions.
A Promptfoo example running an AI Dungeon Master agent with D&D 5e dice, inventory, and stat tools, plus OTLP tracing of its decisions.
A Promptfoo example testing OpenAI's audio-capable models on speech-to-text input and speech-to-speech output.
Test and evaluate OpenAI chat implementations with multi-turn conversation history to ensure proper context preservation.
A promptfoo example for evaluating OpenAI Agent Builder ChatKit workflows via browser automation.
A promptfoo example for evaluating OpenAI's deep research models with web search, code, and MCP tools.
Promptfoo example evaluating OpenAI function calls, comparing external-YAML versus inline function definitions.
A promptfoo example comparing OpenAI's GPT Image models on text-to-image generation, pricing, and quality.
A promptfoo example for evaluating OpenAI's MCP tool integration, including auth and approval workflows.
Promptfoo example testing OpenAI's Realtime API over WebSockets - multi-turn conversation, function calling, and audio input/output.
A promptfoo example collection covering OpenAI Responses API features - reasoning, MCP, caching, and more.
A promptfoo example for testing OpenAI tool-calling accuracy against the Chat Completions API.