Multi-step prompt sequences for complex AI workflows.
19 tools found
Promptfoo example testing Claude Fable 5 on hard coding tasks - debugging, concurrency-aware generation, and security review.
Multi-step prompt workflow that migrates legacy codebases using sandbox agents to isolate shell commands and file edits while keeping orchestration
Promptfoo example that authenticates Anthropic Messages API evals via a local Claude Code OAuth session, not a separate API key.
Promptfoo example wiring Claude to an MCP server via the Anthropic Messages provider, auto-executing tool calls.
A promptfoo example using MLflow AI Gateway as an LLM provider, including as its own grading model.
AI code review CLI from Alibaba that reviews diffs or whole files with line-level precision, using about 1/9 the tokens of general agents.
Promptfoo starting point for comparing Claude and GPT models side by side on your own test cases.
Promptfoo example that compares OpenAI GPT-5.4, Anthropic Claude Sonnet 4.6, and Google Gemini 3.1 Pro Preview on riddle-solving tasks with cost, latency, and
Promptfoo max-score assertion for objective, weighted output selection - deterministic, no extra LLM calls.
Verify LLM outputs against real-time web search results using promptfoo's search-rubric assertion.
A Promptfoo example evaluating tool/function calling across OpenAI, Anthropic, AWS Bedrock, and Groq using a shared weather-lookup function.
Route promptfoo eval traffic through a self-hosted Helicone AI Gateway for unified, load-balanced multi-provider access.