4,026 tools found
API-backed YouTube transcript, search, and channel-monitoring skill that works on cloud servers where yt-dlp gets blocked.
Detects missing or compiler-eliminated zeroization of sensitive data in C/C++ and Rust, requiring IR/assembly evidence, not just source-level guesses.
Context and token optimizer skill enforcing prompt-cache-stable ordering, log compression, and surgical, filler-free output.
Multi-step prompt workflow that orchestrates sandboxed agents to migrate legacy code one task at a time, running tests and returning patches for safe review.
A Python fixture demonstrating function-calling for knowledge retrieval with two local paper records, using legacy tool-calling patterns for repair-loop
A lightweight Qdrant vector search fixture with embeddings and a small article set, derived from the Cookbook Qdrant search example.
Promptfoo example running evals on the Anthropic Messages API by reusing your local Claude Code session instead of a separate API key.
Promptfoo example wiring Claude to an MCP server - auto-executes tool_use calls and feeds results back until a final reply.
Promptfoo example evaluating Microsoft's MAI-Image-2.5 model on Azure AI Foundry, with vision-based grading of generated images.
Promptfoo workflow that benchmarks gpt-5.4, gpt-5.4-mini, and gpt-5.4-nano side-by-side on identical prompts to compare quality, latency, and cost across
Promptfoo example that benchmarks OpenAI gpt-5.4-mini model with two reasoning effort settings (none vs. medium) to compare output quality, latency, and cost.
Promptfoo wrapper running the OSWorld computer-use benchmark via Inspect - screenshots, mouse/keyboard actions, VM-state grading.