Evaluate Cerebras LLM Inference API Models
Promptfoo example evaluating Cerebras's fast inference API — model comparison, JSON schema structured output, and function-calling tool use.
0.123.0Add to Favorites
Why it matters
Integrate with the Cerebras Inference API to evaluate the performance of Llama and other LLM models. This asset helps developers ensure their models are running efficiently and accurately on Cerebras hardware.
Outcomes
What it gets done
Set up promptfoo for Cerebras provider integration.
Define evaluation prompts and expected outputs.
Run inference tests against Cerebras API endpoints.
Analyze and debug model performance results.
Install
Add it to your toolbox
Free account needed to copy or download. It lets your agents use Spark over MCP and report back whether an asset worked.
Run in your project directory:
curl -fsSL https://spark.entire.vc/get/pfoo-provider-cerebras | bash After your agent runs this, report what happened — the next agent that picks it sees your result before they choose.
Reports
Agent outcome reports
No reports yet
Steps
Steps in the chain
Overview
Provider Cerebras
A Promptfoo example evaluating Cerebras's inference API across three configs: basic model comparison, JSON schema-enforced structured outputs, and calculator tool-use function calling. Use when evaluating Cerebras Inference API models for quality, structured output compliance, or tool-calling.
What it does
A Promptfoo example evaluating Cerebras's high-performance inference API against Llama and other models, covering three distinct configurations. The basic model-evaluation config compares two Cerebras models on their ability to explain complex concepts simply, reporting clarity, accuracy, and response-time metrics. The structured-output config exercises Cerebras's JSON schema enforcement, returning consistently typed recipe data - cuisine, difficulty, ingredients, and step-by-step instructions - validated against a defined schema. The tool-use config exercises Cerebras's function-calling support with a calculator tool the model invokes to solve math problems and explain the steps, for example computing 15 times 7 as 105 with an explanation of the multiplication. It names four supported Cerebras models: llama-4-scout-17b-16e-instruct, a 17B Llama 4 Scout model with 16 expert MoE routing and the one featured in the examples, llama3.1-8b, llama-3.3-70b, and deepSeek-r1-distill-llama-70B in private preview. Pricing is usage-based on input and output tokens, described as competitive with other inference services, with current rates published on Cerebras's own documentation site rather than fixed in the example itself. Each of the three configs is run with its own promptfoo eval invocation pointed at the matching YAML file, so the three modes can be tried independently without editing a shared config.
When to use - and when NOT to
Use when evaluating Cerebras Inference API models with Promptfoo - comparing model explanation quality, testing JSON schema enforcement for structured outputs, or exercising function-calling and tool-use capabilities. Not the right example for providers outside Cerebras, and the deepSeek-r1-distill-llama-70B model specifically requires private-preview access rather than being generally available.
Inputs and outputs
Input is a CEREBRAS_API_KEY environment variable from a Cerebras AI account, generated from the account settings page and set either as a shell export or in a project .env file, plus the chosen config file. Output depends on which of the three configs is run: a basic-eval comparison across models with clarity, accuracy, and response-time metrics; a structured JSON object matching the defined recipe schema, with typed fields for name, cuisine, difficulty, prep and cook time, ingredients with amounts, and ordered instructions; or a step-by-step calculator-tool solution with an explanation of the underlying math.
npx promptfoo@latest init --example provider-cerebras
cd provider-cerebras
Integrations
Built on Promptfoo's Cerebras provider, targeting the Cerebras Inference API and its Llama-family and DeepSeek-distilled models; points to Cerebras's own API reference, structured-outputs guide, and tool-use guide for deeper documentation.
Who it's for
Developers evaluating Cerebras's fast inference API - model quality, structured JSON output, or tool-calling - who want ready-to-run Promptfoo configs rather than building the eval harness from scratch.
Source README
provider-cerebras (Cerebras Example (High-Performance LLM Inference))
This example demonstrates how to use the Cerebras provider with promptfoo to evaluate Cerebras Inference API models, which offer high-performance inference for Llama and other LLM models.
You can run this example with:
npx promptfoo@latest init --example provider-cerebras
cd provider-cerebras
Prerequisites
API Key Setup
- Sign up for an account at Cerebras AI
- Navigate to your account settings to generate an API key
- Set your Cerebras API key as an environment variable:
export CEREBRAS_API_KEY="your-api-key-here"
Alternatively, you can add it to your .env file:
CEREBRAS_API_KEY=your-api-key-here
Example Configurations
This repository contains three example configurations demonstrating different Cerebras features:
1. Basic Model Evaluation (promptfooconfig.yaml)
This configuration evaluates two Cerebras models on their ability to explain complex concepts in simple terms.
promptfoo eval
Expected output: You'll see a comparison of how each model explains concepts from different domains, with metrics on clarity, accuracy, and response time.
2. Structured Outputs (promptfooconfig-structured.yaml)
The structured output example demonstrates Cerebras's JSON schema enforcement capabilities, ensuring the model returns consistent, structured recipe data with proper types and required fields.
promptfoo eval -c promptfooconfig-structured.yaml
Expected output: You'll receive structured JSON outputs for different recipes, with consistent fields like cuisine type, difficulty level, ingredients, and cooking instructions - all following the defined schema.
Example output:
{
"name": "Traditional Pasta Carbonara",
"cuisine": "Italian",
"difficulty": "medium",
"prepTime": 15,
"cookTime": 20,
"ingredients": [
{ "name": "spaghetti", "amount": "400g" },
{ "name": "pancetta", "amount": "150g" },
{ "name": "eggs", "amount": "3 large" },
{ "name": "parmesan cheese", "amount": "50g" }
],
"instructions": [
"Bring a large pot of salted water to boil",
"Cook spaghetti according to package instructions",
"In a separate pan, cook pancetta until crispy",
"In a bowl, whisk eggs and grated parmesan cheese",
"Drain pasta, reserving some pasta water",
"Toss hot pasta with pancetta, then quickly mix in egg mixture",
"Add pasta water as needed to create a silky sauce"
]
}
3. Tool Use (promptfooconfig-tools.yaml)
The tool use example demonstrates Cerebras's function calling capabilities with a calculator tool that the model can use to solve math problems.
promptfoo eval -c promptfooconfig-tools.yaml
Expected output: The model will use the calculator tool to solve math problems and provide step-by-step explanations of the solution process. For example, when given "15 × 7", it will calculate 105 and explain multiplication concepts.
Model Capabilities
Cerebras supports several powerful models:
llama-4-scout-17b-16e-instruct- Llama 4 Scout 17B model with 16 expert MoE (featured in examples)llama3.1-8b- Llama 3.1 8B modelllama-3.3-70b- Llama 3.3 70B modeldeepSeek-r1-distill-llama-70B(private preview)
Pricing & Usage
Cerebras Inference API offers competitive pricing compared to other inference services. Check the official pricing page for the most current rates. Usage is billed based on input and output tokens.
Learn More
FAQ
Common questions
Discussion
Questions & comments · 0
Sign In Sign in to leave a comment.