MCP Connector

Evaluate and Improve AI Assistant Performance

MCP server exposing Mandoline's evaluation framework so AI assistants can score, critique, and improve their own responses.

Works with claudecursor

91
Spark score
out of 100
Updated 10 months ago
Version 0.2.0
Models
claudeuniversal

Add to Favorites

Why it matters

Enable AI assistants like Claude and Cursor to critically evaluate and continuously improve their own performance using the Mandoline evaluation framework via the Model Context Protocol.

Outcomes

What it gets done

01

Define custom evaluation metrics for specific tasks.

02

Evaluate prompt/response pairs against defined metrics.

03

Monitor AI assistant performance and identify areas for improvement.

04

Integrate with AI assistants to facilitate self-evaluation.

Install

Add it to your toolbox

Run in your project directory:

curl -fsSL https://spark.entire.vc/get/vb-mandoline | bash

Capabilities

Tools your agent gets

get_server_health

Confirms MCP server availability and returns correct status

create_metric

Defines custom evaluation criteria for your specific tasks

batch_create_metrics

Creates multiple evaluation metrics in a single operation

get_metric

Retrieves details about a specific metric

get_metrics

Views your metrics with filtering and pagination

update_metric

Modifies existing metric definitions

create_evaluation

Evaluates prompt/response pairs against your metrics

batch_create_evaluations

Evaluates the same content across multiple metrics

+3 tools

Overview

Mandoline MCP Server

MCP server that exposes Mandoline's evaluation framework as tools for defining custom metrics and scoring prompt/response pairs, letting AI assistants critique and improve their own outputs. Use the hosted server with an API key from mandoline.ai to add self-evaluation tools to Claude Code, Claude Desktop, Cursor, or Codex; self-host only for local development or contributing.

What it does

Mandoline MCP Server exposes Mandoline's evaluation framework as MCP tools, letting AI assistants such as Claude Code, Claude Desktop, Cursor, and Codex reflect on, critique, and continuously improve their own performance. It provides tools to define custom evaluation metrics, score prompt/response pairs against those metrics individually or in batch, and browse metric and evaluation history, plus a health-check tool and two mirrored documentation resources.

When to use - and when NOT to

Use it when you want an AI assistant to self-evaluate or compare outputs - for example scoring multiple candidate code solutions against custom metrics and picking the best one - using Mandoline's hosted service. This requires a Mandoline account and an API key from mandoline.ai/account. Running the server locally is only needed for local development or contributing to the project; most users should use the hosted server instead of self-hosting.

Capabilities

Health: get_server_health confirms the MCP server is reachable and returning a healthy status payload. Metrics: create_metric defines custom evaluation criteria for specific tasks, batch_create_metrics creates multiple metrics in one operation, get_metric and get_metrics retrieve or browse metrics with filtering and pagination, and update_metric modifies existing definitions. Evaluations: create_evaluation scores prompt/response pairs against metrics, batch_create_evaluations evaluates the same content against multiple metrics at once, get_evaluation and get_evaluations retrieve or browse evaluation history with filtering and pagination, and update_evaluation adds metadata or context to an evaluation. Two resources are also exposed: llms.txt, a Mandoline docs index (tools, tutorials, blogs, leaderboards, SDKs) mirrored from mandoline.ai/llms.txt, and mcp, the MCP setup guide mirrored from mandoline.ai/mcp.

How to install

For Claude Code:

claude mcp add --scope user --transport http mandoline https://mandoline.ai/mcp --header "x-api-key: sk_****"

Replace sk_**** with your API key from mandoline.ai/account; --scope user applies across projects while --scope project scopes to the current one only. Restart active Claude Code sessions afterward and verify with /mcp. Codex, Claude Desktop, and Cursor connect to the same hosted https://mandoline.ai/mcp endpoint with an x-api-key header, using mcp-remote for Codex and Claude Desktop or a direct url/headers block for Cursor. To self-host instead: clone https://github.com/mandoline-ai/mandoline-mcp-server.git, run npm install and npm run build (Node.js 18+ required), then npm start - the server runs on http://localhost:8080 by default, and any client config can point at http://localhost:8080/mcp instead of the hosted URL. Distributed under the Apache-2.0 license.

Who it's for

Developers using Claude Code, Claude Desktop, Cursor, or Codex who want their AI assistant to score, compare, and improve its own responses against custom evaluation metrics rather than relying on unmeasured judgment.

Source README

Mandoline MCP Server

Enable AI assistants like Claude Code, Claude Desktop, and Cursor to reflect on, critique, and continuously improve their own performance using Mandoline's evaluation framework via the Model Context Protocol.


Client Setup

Most users should start here. Use Mandoline's hosted MCP server to integrate evaluation tools into your AI assistant.

For each integration below, replace sk_**** with your actual API key from mandoline.ai/account.

Claude Code

Use the CLI to add the Mandoline MCP server to Claude Code:

claude mcp add --scope user --transport http mandoline https://mandoline.ai/mcp --header "x-api-key: sk_****"

You can use --scope user (across projects) or --scope project (current project only).

Note: Restart any active Claude Code sessions after configuration changes.

Verify: Run /mcp in Claude Code to see Mandoline listed as a connected server:

Tutorial: Watch Claude evaluate multiple code solutions and pick the best one.

Official Documentation: Claude Code MCP Guide

Codex

Use the CLI to add the Mandoline MCP server to Codex:

codex mcp add mandoline --env MANDOLINE_API_KEY=sk_**** -- npx -y mcp-remote https://mandoline.ai/mcp --header 'x-api-key: ${MANDOLINE_API_KEY}'

Note: Restart any active Codex sessions after configuration changes.

Verify: Run /mcp in Codex to see Mandoline listed as a connected server:

Official Documentation: Codex MCP Configuration

Claude Desktop

Edit your configuration file (Settings > Developer > Edit Config):

  • macOS: ~/Library/Application Support/Claude/claude_desktop_config.json
  • Windows: %APPDATA%/Claude/claude_desktop_config.json
{
  "mcpServers": {
    "Mandoline": {
      "command": "npx",
      "args": [
        "-y",
        "mcp-remote",
        "https://mandoline.ai/mcp",
        "--header",
        "x-api-key: ${MANDOLINE_API_KEY}"
      ],
      "env": {
        "MANDOLINE_API_KEY": "sk_****"
      }
    }
  }
}

This configuration applies globally to all conversations.

Note: Restart Claude Desktop after configuration changes.

Verify: Look for Mandoline tools when you click the "Search and tools" button.

Official Documentation: MCP Quickstart Guide

Cursor

Create or edit your MCP configuration file:

{
  "mcpServers": {
    "Mandoline": {
      "url": "https://mandoline.ai/mcp",
      "headers": {
        "x-api-key": "sk_****"
      }
    }
  }
}

You can use your global configuration (affects all projects) ~/.cursor/mcp.json or project-local configuration (current project only) .cursor/mcp.json (in project root)

Note: Restart Cursor after configuration changes.

Verify: Check the Output panel (Ctrl+Shift+U) → "MCP Logs" for successful connection, or look for Mandoline tools in the Composer Agent.

Official Documentation: Cursor MCP Guide


Server Setup

Only needed if you want to run the server locally or contribute to development. Most users should use the hosted server above.

Prerequisites: Node.js 18+ and npm

Installation

  1. Clone and build

    git clone https://github.com/mandoline-ai/mandoline-mcp-server.git
    cd mandoline-mcp-server
    npm install
    npm run build
    
  2. Configure environment (optional)

    cp .env.example .env.local
    # Edit .env.local to customize PORT, LOG_LEVEL, etc.
    
  3. Start the server

    npm start
    

The server runs on http://localhost:8080 by default.

Using Local Server

To use your local server instead of the hosted one, replace https://mandoline.ai/mcp with http://localhost:8080/mcp in the client configurations above.


Usage

Once integrated, you can use Mandoline evaluation tools directly in your AI assistant conversations.

Tools

Health

Tool Purpose
get_server_health Confirm the MCP server is reachable and returning a healthy status payload.

Metrics

Tool Purpose
create_metric Define custom evaluation criteria for your specific tasks
batch_create_metrics Create multiple evaluation metrics in one operation
get_metric Retrieve details about a specific metric
get_metrics Browse your metrics with filtering and pagination
update_metric Modify existing metric definitions

Evaluations

Tool Purpose
create_evaluation Score prompt/response pairs against your metrics
batch_create_evaluations Evaluate the same content against multiple metrics
get_evaluation Retrieve evaluation results and scores
get_evaluations Browse evaluation history with filtering and pagination
update_evaluation Add metadata or context to evaluations

Resources

Resource Description
llms.txt Mandoline docs index (tools, tutorials, blogs, leaderboards, SDKs); mirrored from https://mandoline.ai/llms.txt.
mcp MCP setup guide for assistants; mirrored from https://mandoline.ai/mcp.

FAQ

Common questions

Discussion

Questions & comments · 0

Sign In Sign in to leave a comment.