Multi-step prompt sequences for complex AI workflows.
52 tools found
A fast-executing fixture with two local paper records demonstrating legacy tool-calling patterns, derived from the Cookbook arXiv retrieval example.
Promptfoo example that compares OpenAI gpt-5.4-mini with two reasoning effort levels (none vs. medium) to evaluate output quality, latency, and cost.
Nineteen promptfoo configs for AWS Bedrock - Claude, Nova, Llama, Converse with thinking/MCP, Knowledge Base RAG, and inference profiles.
Promptfoo starting point for comparing Claude and GPT models side by side on your own test cases.
Promptfoo example that compares OpenAI GPT-5.4, Anthropic Claude Sonnet 4.6, and Google Gemini 3.1 Pro Preview on riddle-solving tasks with cost, latency, and
A promptfoo workflow that benchmarks Llama and GPT models side-by-side by running identical prompts through both APIs and comparing their outputs.
Promptfoo workflow that evaluates and compares Mistral and Llama model performance side-by-side using OpenRouter API with customizable prompts and test cases.
Promptfoo example comparing DeepSeek, Mistral, Llama, and Qwen on factual tasks through a single OpenRouter key.
A Promptfoo example comparing GPT-5.4 against GPT-5.4 Mini on riddles and reasoning tasks, checking cost, latency, and answer quality.
Promptfoo starting point for comparing local Ollama-hosted Phi3 and Llama3 models side by side on your own test cases.
Promptfoo example testing the conversation-relevance assertion, which catches a chatbot drifting off-topic mid-conversation.
Promptfoo workflow that evaluates GPT-4o-mini zero-shot sentiment analysis on IMDB reviews using F-score, precision, and recall metrics.