1,210 tools found
Promptfoo example evaluating DeepSeek models on Azure AI Foundry, including reasoning-model config for DeepSeek-R1.
Promptfoo example comparing Llama 4 Maverick and Scout on Azure AI Foundry for code generation quality and speed.
A promptfoo example evaluating Mistral's reasoning, chat, coding, grading, and embedding models with cost-tiered comparisons.
Promptfoo example testing Claude's thinking feature, comparing reasoning quality between the Anthropic API and AWS Bedrock on logic puzzles.
Evaluate code with LLMs and sandboxed execution. Test code quality and correctness automatically.
Promptfoo example covering Google AI Studio function calling, search grounding, code execution, and URL context with Gemini models.
Promptfoo example evaluating audio generation through Google's Live API with Gemini models.
Two promptfoo examples for Ollama: comparing censored vs uncensored local models, and testing function calling.
A promptfoo example for evaluating multi-agent Python openai-agents SDK workflows with trace-level assertions.
A Promptfoo example running an AI Dungeon Master agent with D&D 5e dice, inventory, and stat tools, plus OTLP tracing of its decisions.
A Promptfoo example testing OpenAI's audio-capable models on speech-to-text input and speech-to-speech output.
Promptfoo example evaluating OpenAI function calls, comparing external-YAML versus inline function definitions.