Multi-step prompt sequences for complex AI workflows.
42 tools found
Promptfoo example evaluating Cerebras high-performance inference API - model comparison, structured JSON, tool use.
Minimal promptfoo scaffold for benchmarking Cohere models via a promptfooconfig.yaml eval.
Promptfoo example for writing a custom Python provider that calls OpenAI, with config loaded from external files.
A cookbook guide to building a GPT-5.1 coding agent with the Agents SDK that scaffolds, edits, and iterates on a codebase.
OpenAI Cookbook notebook: run hybrid vector + BM25 search in Weaviate with automatic OpenAI vectorization.
A Promptfoo example evaluating Claude Opus 4.6's coding, bug-diagnosis, and tradeoff-reasoning skills with a 128K extended-thinking budget.
A promptfoo example evaluating Mistral AI's Magistral reasoning and chat models, and using Mistral for LLM-as-judge grading and embeddings.
Evaluate code with LLMs and sandboxed execution. Test code quality and correctness automatically.
Promptfoo examples testing Vertex AI function calling - one validating tool call shape, one actually executing a callback function.
Two promptfoo examples for Ollama: comparing censored vs uncensored local models, and testing function calling.
A promptfoo example for evaluating multi-agent Python openai-agents SDK workflows with trace-level assertions.
A Promptfoo example testing OpenAI's audio-capable models on speech-to-text input and speech-to-speech output.