Multi-step prompt sequences for complex AI workflows.
34 tools found
A notebook that builds a Kangas DataGrid with 2D embedding projections, colored by score and groupable in the browser.
Notebook indexing OpenAI Wikipedia embeddings into Elasticsearch and running kNN semantic search over them.
Promptfoo example evaluating DeepSeek models on Azure AI Foundry, including reasoning-model config for DeepSeek-R1.
Promptfoo example comparing Llama 4 Maverick and Scout on Azure AI Foundry for code generation quality and speed.
A promptfoo example evaluating Mistral AI's Magistral reasoning and chat models, and using Mistral for LLM-as-judge grading and embeddings.
Evaluate code with LLMs and sandboxed execution. Test code quality and correctness automatically.
Configure HTTP provider with TLS/SSL for mutual TLS (mTLS) authentication. Securely connect services using certificate-based security.
Promptfoo config benchmarking model truthfulness against the TruthfulQA dataset, with custom weighted factuality scoring.
A promptfoo example asserting on OpenTelemetry trace spans - counts, durations, and error rates for a RAG agent.
Example configuration for evaluating xAI Grok Voice Agent API conversations using promptfoo's testing framework for real-time voice AI interactions.