3,719 tools found
Nineteen promptfoo configs for AWS Bedrock - Claude, Nova, Llama, Converse with thinking/MCP, Knowledge Base RAG, and inference profiles.
Promptfoo starting point for comparing Claude and GPT models side by side on your own test cases.
Promptfoo eval config benchmarking GPT-5 vs GPT-5 Mini on riddle-solving, with cost/latency/rubric assertions.
Promptfoo MMLU benchmark comparing GPT-5 vs GPT-5 Mini with step-by-step reasoning prompts and format-checking assertions.
A promptfoo workflow that evaluates GPT-4o model behavior across different temperature settings.
Promptfoo example that compares OpenAI GPT-5.4, Anthropic Claude Sonnet 4.6, and Google Gemini 3.1 Pro Preview on riddle-solving tasks with cost, latency, and
Promptfoo starting point for comparing Llama (via Replicate) and GPT models side by side on your own test cases.
Promptfoo starting point for comparing Mistral and Llama models via OpenRouter side by side on your own test cases.
Promptfoo example comparing DeepSeek, Mistral, Llama, and Qwen on factual tasks through a single OpenRouter key.
Promptfoo example comparing GPT-5.4 and GPT-5.4 Mini on riddles with cost, latency, and LLM-graded quality assertions.
Promptfoo starting point for comparing local Ollama-hosted Phi3 and Llama3 models side by side on your own test cases.
Structure and filter AI prompt test cases using metadata columns in CSV/Excel files for organized evaluation workflows.