4,026 tools found
Promptfoo example for evaluating n8n AI agent workflows through a webhook-triggered provider.
Promptfoo example evaluating prompts against Atlas Cloud's deepseek-v3 model, including routing through a custom gateway.
Promptfoo example comparing three Fireworks AI reasoning models on summarization, graded via a Fireworks embedding similarity score.
A promptfoo example using MLflow AI Gateway as an LLM provider, including as its own grading model.
Promptfoo example comparing three NVIDIA NIM-hosted models on a summarization task with deterministic assertions.
Promptfoo example for OrcaRouter, an OpenAI-compatible adaptive routing gateway to multiple upstream models.
Test suite validating promptfoo's own PolicyPlugin redteam generator across single-input, modifier-driven and multi-input generation modes.
Sub-4-bit quantization plus on-policy distillation recovery for any Hugging Face model - one ModelAdapter per family.
Benchmark for evaluating LLM agents on fixing real-world CVEs inside sandboxed Docker containers, scored against maintainer tests.
lowfat is a lightweight CLI that filters verbose shell output and file content before it reaches an AI agent, cutting token cost.
OpenClaw Companion connects a Windows PC to an OpenClaw gateway and lets you choose exactly which Windows capabilities agents can use.
Open Code Review (ocr) is Alibaba's AI code review CLI, trading some recall for high precision at about 1/9 the tokens of a general agent.