Prompt Chain

Evaluate Code Generation with Max Score Selection

Select the best model output objectively using promptfoo's max-score assertion, a weighted alternative to LLM-judged select-best.


92
Spark score
out of 100
Updated 4 months ago
Version 1.0.0
Models

Add to Favorites

Why it matters

Automate the evaluation of code generation prompts by selecting the highest-scoring output based on predefined criteria. This asset helps ensure the quality and correctness of generated code.

Outcomes

What it gets done

01

Define evaluation metrics for code generation.

02

Run multiple code generation prompts.

03

Select the best code output based on evaluation scores.

04

Integrate with code quality assessment tools.

Install

Add it to your toolbox

Run in your project directory:

curl -fsSL https://spark.entire.vc/get/pfoo-eval-max-score-selection | bash

Steps

Steps in the chain

01
Generate Multiple Outputs
02
Evaluate All Assertions
03
Aggregate Scores
04
Select Best Output
05
Display Results

Overview

Eval Max Score Selection

Uses promptfoo's max-score assertion to deterministically select the best of several outputs via a weighted aggregate of other assertion scores, contrasted against the LLM-judged select-best assertion. Use max-score when selection criteria are objective and reproducibility matters; use select-best instead when the criteria are genuinely subjective.

What it does

This example demonstrates the max-score assertion type, which provides a deterministic way to select the best output from multiple providers by aggregating scores from other assertions (correctness, quality, documentation, etc.), applying configurable weights per assertion type, and selecting the output with the highest weighted score. It differs from select-best by being objective (quantifiable scores rather than LLM judgment), deterministic (same inputs always produce the same selection), transparent (a clear weighted-scoring methodology), and cost-effective (no additional LLM calls for selection). Configuration takes a method (average, the default, weighted average of assertion scores, or sum, a weighted sum), a weights map from assertion type to weight (default 1.0 each), and an optional threshold (minimum score required for selection). The selection process runs in five steps: multiple outputs are generated (one per provider), all assertions are evaluated on each output (Python tests verifying correctness as pass=1/fail=0, LLM rubrics scoring quality aspects 0-1, and other assertion types contributing their own scores), max-score calculates each output's weighted score, the output with the highest score is marked as passing, and the results clearly show which output won and why. In the worked example, three outputs with python/documentation/efficiency assertion scores are combined with weights python=3 and llm-rubric=1: Output A (python=1.0, documentation=0.5, efficiency=0.7) scores 0.84, Output B (python=1.0, documentation=0.9, efficiency=0.8) scores 0.94 and is selected, and Output C (python=0.0, documentation=1.0, efficiency=1.0) scores 0.40 despite the highest documentation and efficiency scores, because it fails the heavily-weighted Python correctness test.

When to use - and when NOT to

Use max-score when you have objective criteria such as passing tests or measurable metrics, want reproducible selection results, need to weight different assertion types differently, or want to avoid the extra API cost of an LLM-judged selection.

Use select-best instead when the selection criteria are genuinely subjective or hard to quantify and need nuanced LLM judgment rather than a weighted formula.

Inputs and outputs

Input is promptfooconfig.yaml with a max-score assertion listing method, weights, and optional threshold; requires API keys for OpenAI/Anthropic to generate the candidate outputs (npx promptfoo@latest eval requires those keys). Output is the eval run itself, showing which output won and why based on the weighted scoring breakdown.

Integrations

Aggregates scores from promptfoo's other assertion types (python, javascript, llm-rubric, contains, etc.) rather than integrating with an external service itself; example providers include OpenAI and Anthropic.

Who it's for

Teams comparing multiple provider outputs who want a reproducible, cost-free selection mechanism based on objective, weighted assertion scores.

npx promptfoo@latest init --example eval-max-score-selection
cd eval-max-score-selection
Source README

eval-max-score-selection (Max-Score Selection)

You can run this example with:

npx promptfoo@latest init --example eval-max-score-selection
cd eval-max-score-selection

This example demonstrates the max-score assertion type for objective output selection based on aggregated scores from other assertions.

Overview

The max-score assertion provides a deterministic way to select the best output from multiple providers by:

  • Aggregating scores from other assertions (correctness, quality, documentation, etc.)
  • Applying configurable weights to different assertion types
  • Selecting the output with the highest weighted score
  • Providing objective, reproducible selection criteria

Key Differences from select-best

  • Objective: Uses quantifiable scores rather than LLM judgment
  • Deterministic: Same inputs always produce same selection
  • Transparent: Clear scoring methodology based on weighted assertions
  • Cost-effective: No additional LLM calls for selection

Configuration

- type: max-score
  value:
    method: average # 'average' (default) or 'sum'
    weights:
      python: 3 # Weight for Python code correctness tests
      llm-rubric: 1 # Weight for LLM-evaluated quality rubrics
      javascript: 2 # Weight for JavaScript tests
      contains: 0.5 # Weight for simple string matching
    threshold: 0.7 # Optional minimum score threshold

Options

  • method: How to aggregate scores
    • average (default): Weighted average of assertion scores
    • sum: Weighted sum of assertion scores
  • weights: Map of assertion types to their weights (default: 1.0)
  • threshold: Minimum score required for selection (optional)

Usage

Basic Example

### Run the main example (requires API keys for OpenAI/Anthropic)
npx promptfoo@latest eval

How It Works

  1. Multiple Outputs Generated: Each provider generates a solution
  2. Assertions Evaluated: All assertions run on each output:
    • Python tests verify correctness (pass=1, fail=0)
    • LLM rubrics evaluate quality aspects (0-1 score)
    • Other assertions contribute their scores
  3. Scores Aggregated: Max-score calculates weighted score for each output
  4. Best Selected: Output with highest score is marked as passing
  5. Results Shown: Clear indication of which output won and why

Example Scoring

Given three outputs with these assertion results:

  • Output A: python=1.0, documentation=0.5, efficiency=0.7
  • Output B: python=1.0, documentation=0.9, efficiency=0.8
  • Output C: python=0.0, documentation=1.0, efficiency=1.0

With weights: python=3, llm-rubric=1

  • Output A: (3×1.0 + 1×0.5 + 1×0.7) / 5 = 0.84
  • Output B: (3×1.0 + 1×0.9 + 1×0.8) / 5 = 0.94 ✓ (selected)
  • Output C: (3×0.0 + 1×1.0 + 1×1.0) / 5 = 0.40

When to Use max-score

Use max-score when:

  • You have objective criteria (tests, metrics)
  • You want reproducible results
  • You need to weight different aspects differently
  • You want to avoid additional API costs

Use select-best when:

  • You need subjective judgment
  • The criteria are hard to quantify
  • You want nuanced evaluation of quality

FAQ

Common questions

Discussion

Questions & comments · 0

Sign In Sign in to leave a comment.