3,719 tools found
Context and token optimizer skill enforcing prompt-cache-stable ordering, log compression, and surgical, filler-free output.
Agents SDK cookbook that migrates a legacy codebase shard by shard in isolated sandboxes, while orchestration stays on the trusted host.
A fast-executing fixture with two local paper records demonstrating legacy tool-calling patterns, derived from the Cookbook arXiv retrieval example.
Sampled fixture derived from the Cookbook's Qdrant search example, using a tiny local article set for fast repair-loop validation.
Promptfoo example that authenticates Anthropic Messages API evals via a local Claude Code OAuth session, not a separate API key.
Promptfoo example wiring Claude to an MCP server via the Anthropic Messages provider, auto-executing tool calls.
Promptfoo example evaluating Microsoft MAI image generation on Azure AI Foundry, graded by a vision LLM judge.
Promptfoo example A/B testing two versions of a Codex review skill with max-score for objective winner selection.
Promptfoo workflow that benchmarks gpt-5.4, gpt-5.4-mini, and gpt-5.4-nano side-by-side on identical prompts to compare quality, latency, and cost across
Promptfoo example that compares OpenAI gpt-5.4-mini with two reasoning effort levels (none vs. medium) to evaluate output quality, latency, and cost.
Promptfoo wrapper running the OSWorld computer-use benchmark via Inspect, with reference GPT-5.5 pass-rate data.
Promptfoo example evaluating n8n AI agent workflows via webhook, with session-aware multi-turn support.