4,052 tools found
Promptfoo example for Hugging Face Inference Endpoints with setup instructions and evaluation commands.
Promptfoo example that scores LLM outputs for PII leakage using a Hugging Face classifier you deploy yourself.
Promptfoo example evaluating CrewAI multi-agent performance - Python agent setup, promptfoo provider interface, and reliability caveats.
Docker-based sandbox example for evaluating AI-generated code using promptfoo, available as an initialization template.
Promptfoo workflow that runs LLM prompt tests using test cases stored in Google Sheets, supporting both public and authenticated private sheet access.
Promptfoo example routing evals through a self-hosted Helicone AI Gateway for unified, OpenAI-compatible multi-provider access.
Promptfoo example that runs a Python LangChain LCEL chain and compares it against GPT-5.4 on a math task.
Promptfoo example using Langfuse prompt management with labels - deploy, A/B test and roll back prompt versions per environment without code changes.
Built-in OpenTelemetry tracing for promptfoo LLM evaluations, capturing GenAI semantic conventions and token usage across OpenAI, Anthropic, and 14 supported
The JavaScript OpenTelemetry tracing example for Promptfoo, tracing provider internals and validating tool-call trajectories.
Promptfoo example evaluating a PydanticAI weather agent's structured outputs and tool use via promptfoo's Python provider.
Promptfoo workflow that imports test cases from CSV files stored in Microsoft SharePoint using Azure AD certificate-based authentication.