Evaluate and Improve AI Assistant Performance
MCP server exposing Mandoline's evaluation framework so AI assistants can score, critique, and improve their own responses.
0.2.0Add to Favorites
Why it matters
Enable AI assistants like Claude and Cursor to critically evaluate and continuously improve their own performance using the Mandoline evaluation framework via the Model Context Protocol.
Outcomes
What it gets done
Define custom evaluation metrics for specific tasks.
Evaluate prompt/response pairs against defined metrics.
Monitor AI assistant performance and identify areas for improvement.
Integrate with AI assistants to facilitate self-evaluation.
Install
Add it to your toolbox
Run in your project directory:
curl -fsSL https://spark.entire.vc/get/vb-mandoline | bash Capabilities
Tools your agent gets
Confirms MCP server availability and returns correct status
Defines custom evaluation criteria for your specific tasks
Creates multiple evaluation metrics in a single operation
Retrieves details about a specific metric
Views your metrics with filtering and pagination
Modifies existing metric definitions
Evaluates prompt/response pairs against your metrics
Evaluates the same content across multiple metrics
Overview
Mandoline MCP Server
MCP server that exposes Mandoline's evaluation framework as tools for defining custom metrics and scoring prompt/response pairs, letting AI assistants critique and improve their own outputs. Use the hosted server with an API key from mandoline.ai to add self-evaluation tools to Claude Code, Claude Desktop, Cursor, or Codex; self-host only for local development or contributing.
What it does
Mandoline MCP Server exposes Mandoline's evaluation framework as MCP tools, letting AI assistants such as Claude Code, Claude Desktop, Cursor, and Codex reflect on, critique, and continuously improve their own performance. It provides tools to define custom evaluation metrics, score prompt/response pairs against those metrics individually or in batch, and browse metric and evaluation history, plus a health-check tool and two mirrored documentation resources.
When to use - and when NOT to
Use it when you want an AI assistant to self-evaluate or compare outputs - for example scoring multiple candidate code solutions against custom metrics and picking the best one - using Mandoline's hosted service. This requires a Mandoline account and an API key from mandoline.ai/account. Running the server locally is only needed for local development or contributing to the project; most users should use the hosted server instead of self-hosting.
Capabilities
Health: get_server_health confirms the MCP server is reachable and returning a healthy status payload. Metrics: create_metric defines custom evaluation criteria for specific tasks, batch_create_metrics creates multiple metrics in one operation, get_metric and get_metrics retrieve or browse metrics with filtering and pagination, and update_metric modifies existing definitions. Evaluations: create_evaluation scores prompt/response pairs against metrics, batch_create_evaluations evaluates the same content against multiple metrics at once, get_evaluation and get_evaluations retrieve or browse evaluation history with filtering and pagination, and update_evaluation adds metadata or context to an evaluation. Two resources are also exposed: llms.txt, a Mandoline docs index (tools, tutorials, blogs, leaderboards, SDKs) mirrored from mandoline.ai/llms.txt, and mcp, the MCP setup guide mirrored from mandoline.ai/mcp.
How to install
For Claude Code:
claude mcp add --scope user --transport http mandoline https://mandoline.ai/mcp --header "x-api-key: sk_****"
Replace sk_**** with your API key from mandoline.ai/account; --scope user applies across projects while --scope project scopes to the current one only. Restart active Claude Code sessions afterward and verify with /mcp. Codex, Claude Desktop, and Cursor connect to the same hosted https://mandoline.ai/mcp endpoint with an x-api-key header, using mcp-remote for Codex and Claude Desktop or a direct url/headers block for Cursor. To self-host instead: clone https://github.com/mandoline-ai/mandoline-mcp-server.git, run npm install and npm run build (Node.js 18+ required), then npm start - the server runs on http://localhost:8080 by default, and any client config can point at http://localhost:8080/mcp instead of the hosted URL. Distributed under the Apache-2.0 license.
Who it's for
Developers using Claude Code, Claude Desktop, Cursor, or Codex who want their AI assistant to score, compare, and improve its own responses against custom evaluation metrics rather than relying on unmeasured judgment.
Source README
Mandoline MCP Server
Enable AI assistants like Claude Code, Claude Desktop, and Cursor to reflect on, critique, and continuously improve their own performance using Mandoline's evaluation framework via the Model Context Protocol.
Client Setup
Most users should start here. Use Mandoline's hosted MCP server to integrate evaluation tools into your AI assistant.
For each integration below, replace sk_**** with your actual API key from mandoline.ai/account.
Claude Code
Use the CLI to add the Mandoline MCP server to Claude Code:
claude mcp add --scope user --transport http mandoline https://mandoline.ai/mcp --header "x-api-key: sk_****"
You can use --scope user (across projects) or --scope project (current project only).
Note: Restart any active Claude Code sessions after configuration changes.
Verify: Run /mcp in Claude Code to see Mandoline listed as a connected server:
Tutorial: Watch Claude evaluate multiple code solutions and pick the best one.
Official Documentation: Claude Code MCP Guide
Codex
Use the CLI to add the Mandoline MCP server to Codex:
codex mcp add mandoline --env MANDOLINE_API_KEY=sk_**** -- npx -y mcp-remote https://mandoline.ai/mcp --header 'x-api-key: ${MANDOLINE_API_KEY}'
Note: Restart any active Codex sessions after configuration changes.
Verify: Run /mcp in Codex to see Mandoline listed as a connected server:
Official Documentation: Codex MCP Configuration
Claude Desktop
Edit your configuration file (Settings > Developer > Edit Config):
- macOS:
~/Library/Application Support/Claude/claude_desktop_config.json - Windows:
%APPDATA%/Claude/claude_desktop_config.json
{
"mcpServers": {
"Mandoline": {
"command": "npx",
"args": [
"-y",
"mcp-remote",
"https://mandoline.ai/mcp",
"--header",
"x-api-key: ${MANDOLINE_API_KEY}"
],
"env": {
"MANDOLINE_API_KEY": "sk_****"
}
}
}
}
This configuration applies globally to all conversations.
Note: Restart Claude Desktop after configuration changes.
Verify: Look for Mandoline tools when you click the "Search and tools" button.
Official Documentation: MCP Quickstart Guide
Cursor
Create or edit your MCP configuration file:
{
"mcpServers": {
"Mandoline": {
"url": "https://mandoline.ai/mcp",
"headers": {
"x-api-key": "sk_****"
}
}
}
}
You can use your global configuration (affects all projects) ~/.cursor/mcp.json or project-local configuration (current project only) .cursor/mcp.json (in project root)
Note: Restart Cursor after configuration changes.
Verify: Check the Output panel (Ctrl+Shift+U) → "MCP Logs" for successful connection, or look for Mandoline tools in the Composer Agent.
Official Documentation: Cursor MCP Guide
Server Setup
Only needed if you want to run the server locally or contribute to development. Most users should use the hosted server above.
Prerequisites: Node.js 18+ and npm
Installation
Clone and build
git clone https://github.com/mandoline-ai/mandoline-mcp-server.git cd mandoline-mcp-server npm install npm run buildConfigure environment (optional)
cp .env.example .env.local # Edit .env.local to customize PORT, LOG_LEVEL, etc.Start the server
npm start
The server runs on http://localhost:8080 by default.
Using Local Server
To use your local server instead of the hosted one, replace https://mandoline.ai/mcp with http://localhost:8080/mcp in the client configurations above.
Usage
Once integrated, you can use Mandoline evaluation tools directly in your AI assistant conversations.
Tools
Health
| Tool | Purpose |
|---|---|
get_server_health |
Confirm the MCP server is reachable and returning a healthy status payload. |
Metrics
| Tool | Purpose |
|---|---|
create_metric |
Define custom evaluation criteria for your specific tasks |
batch_create_metrics |
Create multiple evaluation metrics in one operation |
get_metric |
Retrieve details about a specific metric |
get_metrics |
Browse your metrics with filtering and pagination |
update_metric |
Modify existing metric definitions |
Evaluations
| Tool | Purpose |
|---|---|
create_evaluation |
Score prompt/response pairs against your metrics |
batch_create_evaluations |
Evaluate the same content against multiple metrics |
get_evaluation |
Retrieve evaluation results and scores |
get_evaluations |
Browse evaluation history with filtering and pagination |
update_evaluation |
Add metadata or context to evaluations |
Resources
| Resource | Description |
|---|---|
llms.txt |
Mandoline docs index (tools, tutorials, blogs, leaderboards, SDKs); mirrored from https://mandoline.ai/llms.txt. |
mcp |
MCP setup guide for assistants; mirrored from https://mandoline.ai/mcp. |
FAQ
Common questions
Discussion
Questions & comments · 0
Sign In Sign in to leave a comment.