Multi-step prompt sequences for complex AI workflows.
34 tools found
Promptfoo workflow that benchmarks gpt-5.4, gpt-5.4-mini, and gpt-5.4-nano side-by-side on identical prompts to compare quality, latency, and cost across
Promptfoo example that compares OpenAI gpt-5.4-mini with two reasoning effort levels (none vs. medium) to evaluate output quality, latency, and cost.
Promptfoo wrapper running the OSWorld computer-use benchmark via Inspect, with reference GPT-5.5 pass-rate data.
Promptfoo internal suite testing its own PolicyPlugin redteam test generator across five generation modes.
Nineteen promptfoo configs for AWS Bedrock - Claude, Nova, Llama, Converse with thinking/MCP, Knowledge Base RAG, and inference profiles.
A promptfoo workflow that evaluates GPT-4o model behavior across different temperature settings.
Promptfoo example that compares OpenAI GPT-5.4, Anthropic Claude Sonnet 4.6, and Google Gemini 3.1 Pro Preview on riddle-solving tasks with cost, latency, and
Promptfoo workflow that evaluates and compares Mistral and Llama model performance side-by-side using OpenRouter API with customizable prompts and test cases.
Promptfoo example comparing DeepSeek, Mistral, Llama, and Qwen on factual tasks through a single OpenRouter key.
A promptfoo example generating test cases from JavaScript/TypeScript functions, both statically and dynamically.
Promptfoo example workflow for evaluating Markdown rendering with language models.
A promptfoo RAG eval example that intentionally includes unsupported facts to show both passing and failing metrics.