4,026 tools found
Promptfoo example configs for AWS Bedrock model access - Claude, GPT-OSS, Grok, Llama, Nova and the unified Converse API with MCP and RAG.
Bare-bones promptfoo starting example for comparing Anthropic Claude and OpenAI GPT model responses side by side.
A promptfoo example workflow that runs GPT-4o evaluations across multiple temperature settings using the promptfoo eval command.
Promptfoo example that compares OpenAI GPT-5.4, Anthropic Claude Sonnet 4.6, and Google Gemini 3.1 Pro Preview on riddle-solving tasks with cost, latency
A promptfoo workflow that benchmarks Llama and GPT models side-by-side by running identical prompts through both APIs and comparing their outputs.
A promptfoo example that compares Mistral and Llama model outputs using OpenRouter.
Promptfoo example comparing DeepSeek, Mistral, Llama and Qwen on factual accuracy through a single OpenRouter API key.
Promptfoo example comparing gpt-5.6-luna, gpt-5.6-terra, gpt-5.6-sol, and gpt-6-astra on riddles at low reasoning effort, side by side.
Bare-bones promptfoo example comparing Microsoft Phi3 and Meta Llama3 locally via Ollama - no API keys or cloud cost required.
Promptfoo example demonstrating CSV and Excel test files with metadata columns for organizing, categorizing, and filtering AI prompt evaluation test cases.
Promptfoo example covering every prompt format it supports - text, file-based, Jinja2, and JS/TS/Python prompt functions returning dynamic config.
Promptfoo example that generates variable values at runtime with a Python or JavaScript script.