Multi-step prompt sequences for complex AI workflows.
42 tools found
Promptfoo example comparing Moonshot's kimi-k3 and kimi-k2.6 models on a summarization task with plain keyword assertions.
Promptfoo example testing Claude Fable 5 on hard coding tasks - debugging, concurrency-aware generation, and security review.
Multi-step prompt workflow that migrates legacy codebases using sandbox agents to isolate shell commands and file edits while keeping orchestration
Promptfoo example that authenticates Anthropic Messages API evals via a local Claude Code OAuth session, not a separate API key.
Promptfoo example evaluating Microsoft's MAI-Image-2.5 model on Azure AI Foundry, with vision-based grading of generated images.
A promptfoo example that scores two Codex skill versions against the same review tasks and auto-selects the stronger output.
Promptfoo workflow that benchmarks gpt-5.4, gpt-5.4-mini, and gpt-5.4-nano side-by-side on identical prompts to compare quality, latency, and cost across
Promptfoo example comparing three Fireworks AI serverless chat models on summarization, graded via similarity.
A promptfoo example using MLflow AI Gateway as an LLM provider, including as its own grading model.
Compares Llama 3.3, Nemotron 70B, and Qwen 2.5 Coder models hosted on NVIDIA NIM in a promptfoo eval.
Promptfoo example for OrcaRouter, an OpenAI-compatible adaptive routing gateway to multiple upstream models.
Promptfoo internal suite testing its own PolicyPlugin redteam test generator across five generation modes.