3,603 tools found
A promptfoo example for evaluating Azure AI Foundry agents through the newer v2 Responses runtime.
Promptfoo example comparing Llama 4 Maverick and Scout on Azure AI Foundry for code generation quality and speed.
Promptfoo example evaluating Mistral's chat, reasoning, grading, and embedding models side by side, from Magistral to Mistral Small 4.
A promptfoo example for evaluating Azure OpenAI text generation and vision models, including three image input methods.
Promptfoo eval config for testing AWS Bedrock Agents, including conversation memory retention across a session.
Promptfoo example testing Claude's thinking feature, comparing reasoning quality between the Anthropic API and AWS Bedrock on logic puzzles.
A promptfoo example routing OpenAI, Anthropic, and Groq requests through Cloudflare AI Gateway.
Evaluate code with LLMs and sandboxed execution. Test code quality and correctness automatically.
Promptfoo config testing ElevenLabs conversational AI agents on weighted criteria, tool use, cost, and latency across multi-turn conversations.
A promptfoo example aligning audio to transcripts with ElevenLabs Forced Alignment, outputting JSON or SRT subtitles.
Promptfoo config testing ElevenLabs Audio Isolation, verifying background-noise removal across formats, quality settings, and speaker counts.
Test ElevenLabs STT for audio transcription accuracy. This promptfoo example helps evaluate transcription quality.