3,719 tools found
A promptfoo example demonstrating G-Eval usage, referencing the research paper arxiv.org/abs/2303.16634.
Promptfoo example evaluating GPT-4o vision models on Fashion MNIST image classification with a structured schema.
Run external JavaScript assertions to validate AI-generated code quality and correctness with automated testing workflows.
A promptfoo example for evaluating and validating JSON-formatted model outputs.
A promptfoo example for evaluating how a model's markdown-formatted output renders.
Promptfoo max-score assertion for objective, weighted output selection - deterministic, no extra LLM calls.
Create dynamic, static, and derived custom metric names in promptfoo assertions for more filterable eval results.
A promptfoo RAG eval example that intentionally includes unsupported facts to show both passing and failing metrics.
A full promptfoo RAG example over SEC filings, with a PDF ingest pipeline and a Python retrieval provider.
Verify LLM outputs against real-time web search results using promptfoo's search-rubric assertion.
Evaluate multiple AI outputs side-by-side and automatically select the best response based on your quality criteria.
A promptfoo example that demonstrates how to have an LLM grade its own output according to predefined expectations, with configuration via YAML or CSV test