Multi-step prompt sequences for complex AI workflows.
12 tools found
Compares Llama 3.3, Nemotron 70B, and Qwen 2.5 Coder models hosted on NVIDIA NIM in a promptfoo eval.
A promptfoo workflow that benchmarks Llama and GPT models side-by-side by running identical prompts through both APIs and comparing their outputs.
Promptfoo workflow that evaluates and compares Mistral and Llama model performance side-by-side using OpenRouter API with customizable prompts and test cases.
Promptfoo example comparing DeepSeek, Mistral, Llama, and Qwen on factual tasks through a single OpenRouter key.
Promptfoo starting point for comparing local Ollama-hosted Phi3 and Llama3 models side by side on your own test cases.
Promptfoo example evaluating Cerebras high-performance inference API - model comparison, structured JSON, tool use.
Compares IBM Granite, Meta Llama, and Mistral models available through IBM watsonx.ai in promptfoo.
Notebook using LlamaIndex to extract and compare insights across long SEC 10-K filings with a RAG pipeline.
Promptfoo example comparing GPT, Claude, Llama, and Mistral models side by side on Azure AI Foundry.
Promptfoo example comparing Llama 4 Maverick and Scout on Azure AI Foundry for code generation quality and speed.
Two promptfoo examples for Ollama: comparing censored vs uncensored local models, and testing function calling.
A promptfoo example moderating GPT-5-mini output with LlamaGuard 3 and 4 via Replicate, including category-specific checks.