Trace LLM Operations with OpenTelemetry
A promptfoo example asserting on OpenTelemetry trace spans from a Python provider using protobuf-format OTLP export.
Maintainer of this project? Claim this page to edit the listing.
code-scan-action-0.1Add to Favorites
Why it matters
Integrate OpenTelemetry tracing into your Python applications to gain visibility into LLM provider operations during Promptfoo evaluations. Understand and debug internal processes efficiently.
Outcomes
What it gets done
Instrument Python code for OpenTelemetry tracing.
Export trace data using the efficient protobuf format.
Analyze LLM provider internal operations during evaluations.
Debug and optimize performance with detailed trace information.
Install
Add it to your toolbox
Run in your project directory:
curl -fsSL https://spark.entire.vc/get/pfoo-opentelemetry-tracing-python | bash Overview
Opentelemetry Tracing Python
This promptfoo example asserts on OpenTelemetry trace spans from a Python-based RAG agent provider, accepting protobuf-format OTLP data in addition to JSON. Use it when the traced provider is Python-based and exports OTLP data in protobuf format; it requires the provider to emit the expected span names.
What it does
This promptfoo example is the Python counterpart to promptfoo's OpenTelemetry tracing setup: a python:provider.py provider emits OTLP traces for a simulated RAG agent workflow, and tests assert on span counts, durations, and error rates. It specifically accepts protobuf-format OTLP data (Python's OTLP exporter uses protobuf by default), in addition to JSON.
When to use - and when NOT to
Use this example when your traced provider is written in Python and its OpenTelemetry exporter sends protobuf-encoded OTLP data rather than JSON - the tracing config here explicitly lists acceptFormats: ['json', 'protobuf'] to handle that. It requires the provider under test to emit spans matching the expected names (rag_agent_workflow, retrieve_document_*, reasoning_*, llm_generation); it is not useful for a provider that doesn't produce compatible trace spans.
Inputs and outputs
Each test enables tracing via metadata: { tracingEnabled: true } and asserts on expected spans, for example:
assert:
# Ensure the main workflow span exists
- type: trace-span-count
value:
pattern: 'rag_agent_workflow'
min: 1
max: 1
# Ensure the LLM generation span exists
- type: trace-span-count
value:
pattern: 'llm_generation'
min: 1
# Ensure the overall workflow completes quickly
- type: trace-span-duration
value:
pattern: 'rag_agent_workflow'
max: 5000 # 5 seconds max
defaultTest adds a metric-only p95 latency check (weight: 0) and a hard cap of 500ms per document-retrieval span. Tracing is configured to accept both formats:
tracing:
enabled: true
otlp:
http:
enabled: true
port: 4318
# Python's OTLP exporter uses protobuf by default
acceptFormats: ['json', 'protobuf']
Integrations
Runs against a Python provider (python:provider.py) that emits OTLP traces, received by promptfoo's built-in OTLP HTTP endpoint (port 4318) in either JSON or protobuf format, using the trace-span-count, trace-span-duration, and trace-error-spans assertion types.
Who it's for
Teams building Python-based, OpenTelemetry-instrumented LLM agents who need promptfoo to validate internal execution behavior (spans, durations, error rates) via protobuf-format OTLP traces.
Source README
yaml-language-server: $schema=https://promptfoo.dev/config-schema.json
description: OpenTelemetry tracing with Python (protobuf format)
providers:
- python:provider.py
prompts:
- 'Explain how {{topic}} works in simple terms'
tests:
vars:
topic: 'quantum computing'
metadata:
tracingEnabled: true
testCaseId: 'python-test-1'
assert:Ensure the main workflow span exists
- type: trace-span-count
value:
pattern: 'rag_agent_workflow'
min: 1
max: 1
Ensure we retrieve exactly 3 documents
- type: trace-span-count
value:
pattern: 'retrieve_document_*'
min: 3
max: 3
Ensure all reasoning steps occur
- type: trace-span-count
value:
pattern: 'reasoning_*'
min: 3
Ensure the LLM generation span exists
- type: trace-span-count
value:
pattern: 'llm_generation'
min: 1
Ensure the overall workflow completes quickly
- type: trace-span-duration
value:
pattern: 'rag_agent_workflow'
max: 5000 # 5 seconds max
Ensure no errors occur
- type: trace-error-spans
value:
max_count: 0
- type: trace-span-count
vars:
topic: 'machine learning'
metadata:
tracingEnabled: true
testCaseId: 'python-test-2'
assert:type: trace-span-count
value:
pattern: 'rag_agent_workflow'
min: 1
max: 1type: trace-span-count
value:
pattern: 'retrieve_document_*'
min: 3
max: 3type: trace-span-count
value:
pattern: 'reasoning_*'
min: 3type: trace-span-duration
value:
pattern: 'rag_agent_workflow'
max: 5000type: trace-error-spans
value:
max_count: 0
Default assertions for all test cases
defaultTest:
assert:
# Monitor overall latency
- type: trace-span-duration
value:
pattern: '*'
max: 2000
percentile: 95
weight: 0
metric: p95_latency
# Ensure retrieval operations are fast
- type: trace-span-duration
value:
pattern: 'retrieve_document_*'
max: 500
Tracing configuration - note we accept both JSON and protobuf
tracing:
enabled: true
otlp:
http:
enabled: true
port: 4318
# Python's OTLP exporter uses protobuf by default
acceptFormats: ['json', 'protobuf']
FAQ
Common questions
Discussion
Questions & comments ยท 0
Sign In Sign in to leave a comment.