Multi-step prompt sequences for complex AI workflows.
52 tools found
Notebook using LlamaIndex to extract and compare insights across long SEC 10-K filings with a RAG pipeline.
OpenAI cookbook comparing naive RAG, distance filtering, and HyDE for robust question answering with Chroma.
Walkthrough for setting up Weaviate with OpenAI's vectorizer and Q&A modules to answer questions over your own data.
Example workflows for evaluating Amazon Bedrock models, agents, and video generation using promptfoo, including Claude, Llama, Mistral, Nova, and tool use.
A Promptfoo example demonstrating Anthropic's web search and web fetch tools with domain filtering and source citations.
A Promptfoo example evaluating Claude Opus, Sonnet, and Haiku models deployed on Azure AI Foundry on explanation tasks.
Promptfoo example comparing GPT, Claude, Llama, and Mistral models side by side on Azure AI Foundry.
Promptfoo example evaluating DeepSeek models on Azure AI Foundry, including reasoning-model config for DeepSeek-R1.
Promptfoo example comparing Llama 4 Maverick and Scout on Azure AI Foundry for code generation quality and speed.
Promptfoo example testing Claude's thinking feature, comparing reasoning quality between the Anthropic API and AWS Bedrock on logic puzzles.
Promptfoo config benchmarking model truthfulness against the TruthfulQA dataset, with custom weighted factuality scoring.
A promptfoo example for evaluating OpenAI's MCP tool integration, including auth and approval workflows.