Prompt Chain

Trace LLM Operations with OpenTelemetry

A promptfoo example asserting on OpenTelemetry trace spans from a Python provider using protobuf-format OTLP export.

Works with opentelemetrypython

Maintainer of this project? Claim this page to edit the listing.


91
Spark score
out of 100
Updated last month
Version code-scan-action-0.1

Add to Favorites

Why it matters

Integrate OpenTelemetry tracing into your Python applications to gain visibility into LLM provider operations during Promptfoo evaluations. Understand and debug internal processes efficiently.

Outcomes

What it gets done

01

Instrument Python code for OpenTelemetry tracing.

02

Export trace data using the efficient protobuf format.

03

Analyze LLM provider internal operations during evaluations.

04

Debug and optimize performance with detailed trace information.

Install

Add it to your toolbox

Run in your project directory:

curl -fsSL https://spark.entire.vc/get/pfoo-opentelemetry-tracing-python | bash

Overview

Opentelemetry Tracing Python

This promptfoo example asserts on OpenTelemetry trace spans from a Python-based RAG agent provider, accepting protobuf-format OTLP data in addition to JSON. Use it when the traced provider is Python-based and exports OTLP data in protobuf format; it requires the provider to emit the expected span names.

What it does

This promptfoo example is the Python counterpart to promptfoo's OpenTelemetry tracing setup: a python:provider.py provider emits OTLP traces for a simulated RAG agent workflow, and tests assert on span counts, durations, and error rates. It specifically accepts protobuf-format OTLP data (Python's OTLP exporter uses protobuf by default), in addition to JSON.

When to use - and when NOT to

Use this example when your traced provider is written in Python and its OpenTelemetry exporter sends protobuf-encoded OTLP data rather than JSON - the tracing config here explicitly lists acceptFormats: ['json', 'protobuf'] to handle that. It requires the provider under test to emit spans matching the expected names (rag_agent_workflow, retrieve_document_*, reasoning_*, llm_generation); it is not useful for a provider that doesn't produce compatible trace spans.

Inputs and outputs

Each test enables tracing via metadata: { tracingEnabled: true } and asserts on expected spans, for example:

    assert:
      # Ensure the main workflow span exists
      - type: trace-span-count
        value:
          pattern: 'rag_agent_workflow'
          min: 1
          max: 1

      # Ensure the LLM generation span exists
      - type: trace-span-count
        value:
          pattern: 'llm_generation'
          min: 1

      # Ensure the overall workflow completes quickly
      - type: trace-span-duration
        value:
          pattern: 'rag_agent_workflow'
          max: 5000 # 5 seconds max

defaultTest adds a metric-only p95 latency check (weight: 0) and a hard cap of 500ms per document-retrieval span. Tracing is configured to accept both formats:

tracing:
  enabled: true
  otlp:
    http:
      enabled: true
      port: 4318
      # Python's OTLP exporter uses protobuf by default
      acceptFormats: ['json', 'protobuf']

Integrations

Runs against a Python provider (python:provider.py) that emits OTLP traces, received by promptfoo's built-in OTLP HTTP endpoint (port 4318) in either JSON or protobuf format, using the trace-span-count, trace-span-duration, and trace-error-spans assertion types.

Who it's for

Teams building Python-based, OpenTelemetry-instrumented LLM agents who need promptfoo to validate internal execution behavior (spans, durations, error rates) via protobuf-format OTLP traces.

Source README

yaml-language-server: $schema=https://promptfoo.dev/config-schema.json

description: OpenTelemetry tracing with Python (protobuf format)

providers:

  • python:provider.py

prompts:

  • 'Explain how {{topic}} works in simple terms'

tests:

  • vars:
    topic: 'quantum computing'
    metadata:
    tracingEnabled: true
    testCaseId: 'python-test-1'
    assert:

    Ensure the main workflow span exists

    • type: trace-span-count
      value:
      pattern: 'rag_agent_workflow'
      min: 1
      max: 1

    Ensure we retrieve exactly 3 documents

    • type: trace-span-count
      value:
      pattern: 'retrieve_document_*'
      min: 3
      max: 3

    Ensure all reasoning steps occur

    • type: trace-span-count
      value:
      pattern: 'reasoning_*'
      min: 3

    Ensure the LLM generation span exists

    • type: trace-span-count
      value:
      pattern: 'llm_generation'
      min: 1

    Ensure the overall workflow completes quickly

    • type: trace-span-duration
      value:
      pattern: 'rag_agent_workflow'
      max: 5000 # 5 seconds max

    Ensure no errors occur

    • type: trace-error-spans
      value:
      max_count: 0
  • vars:
    topic: 'machine learning'
    metadata:
    tracingEnabled: true
    testCaseId: 'python-test-2'
    assert:

    • type: trace-span-count
      value:
      pattern: 'rag_agent_workflow'
      min: 1
      max: 1

    • type: trace-span-count
      value:
      pattern: 'retrieve_document_*'
      min: 3
      max: 3

    • type: trace-span-count
      value:
      pattern: 'reasoning_*'
      min: 3

    • type: trace-span-duration
      value:
      pattern: 'rag_agent_workflow'
      max: 5000

    • type: trace-error-spans
      value:
      max_count: 0

Default assertions for all test cases

defaultTest:
assert:
# Monitor overall latency
- type: trace-span-duration
value:
pattern: '*'
max: 2000
percentile: 95
weight: 0
metric: p95_latency

# Ensure retrieval operations are fast
- type: trace-span-duration
  value:
    pattern: 'retrieve_document_*'
    max: 500

Tracing configuration - note we accept both JSON and protobuf

tracing:
enabled: true
otlp:
http:
enabled: true
port: 4318
# Python's OTLP exporter uses protobuf by default
acceptFormats: ['json', 'protobuf']

FAQ

Common questions

Discussion

Questions & comments ยท 0

Sign In Sign in to leave a comment.