4,052 tools found
Promptfoo example comparing OpenAI's Sora video models (sora-2 vs sora-2-pro) on quality and cost, with generated videos viewable in promptfoo's UI.
Promptfoo example asserting directly on OpenTelemetry traces from a simulated RAG agent - span counts, durations and error rates, not just text output.
A Promptfoo example tracing a Python LLM provider's internals via OpenTelemetry, exporting spans in protobuf format for span-level eval assertions.
Promptfoo example comparing DeepSeek R1, GPT-4.1 Mini and Claude 4 Sonnet on a joke-telling task through AI/ML API's single-key access to 300+ models.
Promptfoo example running both a standard quality eval and an automated adversarial red-team pass against DeepSeek R1 via OpenRouter.
This promptfoo example evaluates Replicate's Llama 3 text generation and Stable Diffusion XL image generation in one config.
Promptfoo config comparing five Replicate image models - FLUX variants and Stable Diffusion XL - across photorealistic, artistic and product genres.
Promptfoo example comparing GPT-4.1-mini against a Replicate-hosted Llama 2 model reached through an OpenAI-compatible proxy endpoint.
A promptfoo example moderating GPT-5-mini output with LlamaGuard 3 and 4 via Replicate, including category-specific checks.
Promptfoo example running Meta's Llama 4 Scout via Replicate, testing reasoning and creative tasks against Llama 3.
Promptfoo example demonstrating CSV and Excel files as test-case sources, with JSON-valued fields, illustrated through a translation-style eval.
Simple MCP is a Promptfoo example workflow for evaluating MCP servers through direct tool calling tests, security vulnerability checks, and edge case