3,719 tools found
Configures Atlas Cloud's OpenAI-compatible LLM API in promptfoo for multi-model-family evaluation.
Promptfoo example comparing three Fireworks AI serverless chat models on summarization, graded via similarity.
A promptfoo example using MLflow AI Gateway as an LLM provider, including as its own grading model.
Compares Llama 3.3, Nemotron 70B, and Qwen 2.5 Coder models hosted on NVIDIA NIM in a promptfoo eval.
Promptfoo example calling a specific model or OrcaRouter's adaptive auto-router through one OpenAI-compatible endpoint.
Promptfoo internal suite testing its own PolicyPlugin redteam test generator across five generation modes.
Sub-4-bit quantization plus on-policy distillation recovery for any Hugging Face model - one ModelAdapter per family.
Benchmark for evaluating LLM agents on fixing real-world CVEs inside sandboxed Docker containers, scored against maintainer tests.
Lightweight CLI that filters verbose command output and file content before it reaches your AI agent, cutting token usage.
Native Windows tray companion for OpenClaw with quick-send hotkey, web chat, and Node Mode letting the agent control your PC.
AI code review CLI from Alibaba that reviews diffs or whole files with line-level precision, using about 1/9 the tokens of general agents.
Formally verified multipolygon intersection algorithm proven correct in Lean 4, with the implementation written autonomously by AI agents.