4,026 tools found
Promptfoo example for Azure AI Foundry's v2 Responses agent runtime, evaluating an already-defined Foundry agent via the newer @azure/ai-projects SDK.
Promptfoo example comparing Llama 4 Maverick and Scout (plus older Llama 3.x models) on Azure AI Foundry for code-generation quality versus speed.
Compare Mistral AI's Magistral reasoning, chat, and multimodal models in promptfoo, including Mistral-powered grading and embeddings.
Promptfoo example evaluating Azure OpenAI deployments for both text generation and vision capabilities, with separate configs for each.
Promptfoo config for evaluating AWS Bedrock Agents, covering single-agent and supervisor-routed multi-agent setups.
Promptfoo example comparing Claude's extended-thinking reasoning quality between Sonnet 4 via the Anthropic API and Haiku 4.5 via AWS Bedrock.
Promptfoo example routing OpenAI, Anthropic, and Groq requests through Cloudflare AI Gateway for caching and analytics.
E2B Code Eval is a promptfoo workflow that generates Python functions with LLMs, executes them in sandboxed E2B environments, and produces JSON metrics
Promptfoo example testing ElevenLabs Conversational AI voice agents - multi-turn dialogue scored on weighted criteria, tool use, cost, and latency.
A promptfoo example that evaluates ElevenLabs forced alignment, generating time-aligned SRT, VTT, or JSON subtitles from audio and transcript.
Promptfoo config testing ElevenLabs' Audio Isolation API - noise removal across seven input formats, three output formats, and cost tracking.
Promptfoo example that tests ElevenLabs speech-to-text with WER scoring, speaker diarization, and cost and latency checks.