3,719 tools found
Validate that model-generated SQL is syntactically correct using node-sql-parser in a promptfoo eval.
Evaluate function/tool calling across OpenAI, Anthropic, AWS Bedrock, and Groq in a single promptfoo eval.
Evaluate LLM factuality on the TruthfulQA dataset from HuggingFace, scoring five distinct factuality categories.
Automatically detect and classify hate speech in text using Hugging Face models to moderate content and protect online communities.
Set up and evaluate Hugging Face inference endpoints with automated testing workflows and configuration templates.
Automatically detect and extract personally identifiable information from text using Hugging Face models for privacy compliance.
Test web applications with Playwright browser automation as a promptfoo provider, against a local Gradio demo app.
Evaluate CrewAI multi-agent orchestration performance using promptfoo, including reliability caveats for complex queries.
Test AI-generated code safely in isolated Docker sandboxes before deployment, validating outputs in controlled environments.
A promptfoo pipeline that has an LLM write code, runs it in an e2b sandbox, and grades it with generated unit tests.
Connect Google Sheets to AI testing workflows for spreadsheet-based test case management and automated result reporting.
Route promptfoo eval traffic through a self-hosted Helicone AI Gateway for unified, load-balanced multi-provider access.