Test Audio Transcription with ElevenLabs
Promptfoo example that tests ElevenLabs speech-to-text with WER scoring, speaker diarization, and cost and latency checks.
0.123.0Add to Favorites
Why it matters
Automate the testing of audio transcription services using the ElevenLabs STT provider. Ensure the accuracy and reliability of your speech-to-text capabilities.
Outcomes
What it gets done
Configure ElevenLabs STT for transcription testing.
Process audio files for transcription.
Evaluate the quality of transcribed text.
Integrate transcription testing into CI/CD pipelines.
Install
Add it to your toolbox
Free account needed to copy or download. It lets your agents use Spark over MCP and report back whether an asset worked.
Run in your project directory:
curl -fsSL https://spark.entire.vc/get/pfoo-elevenlabs-stt | bash After your agent runs this, report what happened — the next agent that picks it sees your result before they choose.
Reports
Agent outcome reports
No reports yet
Overview
Elevenlabs Stt
This promptfoo example runs audio through ElevenLabs' Speech-to-Text API and scores the result with Word Error Rate, speaker diarization, and cost and latency assertions, using config-, prompt-, or vars-level audio input. Use it to regression-test an ElevenLabs STT integration's accuracy, diarization, cost, and latency; it requires a real ElevenLabs API key and account.
What it does
This is a promptfoo example (provider-elevenlabs/stt) for testing ElevenLabs' Speech-to-Text API as a promptfoo provider. It runs audio files through the eleven_speech_to_text_v1 model and lets you assert on the results: automatic Word Error Rate (WER) calculation against a reference transcript, with substitutions, deletions, insertions, and a per-word breakdown; speaker diarization, labeling and timestamping up to a configurable number of distinct speakers; cost and latency thresholds; and content assertions (contains/not-contains on the transcribed text, or custom JavaScript assertions on WER or speaker count). Audio can be supplied at the config, prompt, or test-vars level, and the provider supports MP3, WAV, FLAC, M4A, OGG, Opus, and WebM.
When to use - and when NOT to
Use it as a template when you need to benchmark or regression-test an STT integration: comparing transcription accuracy across audio qualities or languages, verifying diarization correctly separates speakers, or gating a build on cost, latency, or WER thresholds. It needs an ELEVENLABS_API_KEY and real ElevenLabs usage - the free tier covers 1 hour/month, paid tiers run about $0.10 per minute - so it's not a fit for testing without an ElevenLabs account, and WER comparisons are only meaningful when the reference text exactly matches the audio, including punctuation.
Inputs and outputs
Inputs: one or more audio files, via audioFile config, a prompt list, or vars, plus config like modelId, language (ISO 639-1, or omitted for auto-detect), diarization/maxSpeakers, and referenceText/calculateWER. Output is the transcription text plus, when enabled, a diarization array (per-speaker segments with start and end times and confidence) and a wer metrics object (error rate, substitution/deletion/insertion counts, and an aligned reference-vs-hypothesis diff) available to test assertions via context.vars.metadata.
Integrations
- promptfoo eval framework, run via
npx promptfoo@latest eval - ElevenLabs Speech-to-Text API (
eleven_speech_to_text_v1), 30+ languages - Cost and latency assertion types built into promptfoo
- Related promptfoo examples: ElevenLabs TTS and ElevenLabs Isolation for audio cleanup
- Config options cover
baseUrl, requesttimeout(default 120000ms), andretries(default 3), so the provider can point at a non-default endpoint or tolerate flaky connections
Troubleshooting notes in the example cover a missing ELEVENLABS_API_KEY, unsupported audio formats (convert with ffmpeg), and unexpectedly high WER on clear audio - usually a reference-text mismatch, wrong auto-detected language, or background noise, since the WER calculation normalizes case and punctuation before comparing.
Who it's for
Teams building or evaluating a voice product who want a promptfoo-based way to regression-test ElevenLabs transcription accuracy, speaker diarization, cost, and latency before shipping.
Source README
provider-elevenlabs/stt (ElevenLabs Speech-to-Text)
This example demonstrates how to use ElevenLabs STT provider for audio transcription testing.
Quick Start
npx promptfoo@latest init --example provider-elevenlabs/stt
cd provider-elevenlabs/stt
export ELEVENLABS_API_KEY=your_api_key_here
npx promptfoo@latest eval
Features
- Audio Transcription: Convert speech to text with high accuracy
- Speaker Diarization: Identify and separate multiple speakers in audio
- Word Error Rate (WER): Measure transcription accuracy against reference text
- Multi-format Support: MP3, WAV, FLAC, M4A, OGG, OPUS, WebM
Setup
Set your API key:
export ELEVENLABS_API_KEY=your_api_key_herePrepare audio files:
Create anaudio/directory with your test audio files:mkdir -p audio # Place your audio files in the audio/ directoryRun the evaluation:
promptfoo eval
Configuration
Basic Transcription
providers:
- id: elevenlabs:stt:basic
config:
modelId: eleven_speech_to_text_v1
language: en # ISO 639-1 language code
Speaker Diarization
Identify and label different speakers in your audio:
providers:
- id: elevenlabs:stt:diarization
config:
modelId: eleven_speech_to_text_v1
diarization: true
maxSpeakers: 3 # Optional: hint for expected number of speakers
The response will include speaker segments:
{
"text": "Full transcription...",
"diarization": [
{
"speaker_id": "speaker_0",
"text": "Hello, how are you?",
"start_time_ms": 0,
"end_time_ms": 2500,
"confidence": 0.95
},
{
"speaker_id": "speaker_1",
"text": "I'm doing well, thanks!",
"start_time_ms": 2500,
"end_time_ms": 5000,
"confidence": 0.92
}
]
}
Accuracy Testing with WER
Word Error Rate (WER) measures transcription accuracy. Lower is better (0 = perfect).
providers:
- id: elevenlabs:stt:accuracy
config:
modelId: eleven_speech_to_text_v1
calculateWER: true
referenceText: The quick brown fox jumps over the lazy dog
WER Formula: (Substitutions + Deletions + Insertions) / Total Words
The response includes detailed WER metrics:
{
"wer": 0.05, // 5% error rate
"substitutions": 1,
"deletions": 0,
"insertions": 0,
"correct": 19,
"totalWords": 20,
"details": {
"reference": "the quick brown fox jumps",
"hypothesis": "the quick green fox jumps",
"alignment": "REF: the quick brown fox jumps\nHYP: the quick green fox jumps\nOPS: SSSSS"
}
}
WER Interpretation:
- 0.00 - 0.05: Excellent (95%+ accurate)
- 0.05 - 0.10: Good (90-95% accurate)
- 0.10 - 0.20: Fair (80-90% accurate)
- 0.20+: Poor (< 80% accurate)
Supported Audio Formats
| Format | Extension | Notes |
|---|---|---|
| MP3 | .mp3 | Widely compatible |
| MP4 Audio | .mp4, .m4a | AAC/MPEG-4 audio |
| WAV | .wav | Uncompressed, high quality |
| FLAC | .flac | Lossless compression |
| OGG | .ogg | Open format |
| Opus | .opus | Modern, efficient codec |
| WebM | .webm | Web-optimized |
Audio Input Methods
Method 1: Config-level
providers:
- id: elevenlabs:stt
config:
audioFile: path/to/audio.mp3
Method 2: Prompt-level
prompts:
- audio/sample1.mp3
- audio/sample2.wav
Method 3: Vars-level
tests:
- vars:
audioFile: audio/sample.mp3
Testing Assertions
Cost Threshold
tests:
- assert:
- type: cost
threshold: 0.05 # Max $0.05 per transcription
Latency Threshold
tests:
- assert:
- type: latency
threshold: 10000 # Max 10 seconds
Transcription Quality
tests:
- assert:
- type: contains
value: expected phrase
- type: not-contains
value: incorrect phrase
WER Threshold
tests:
- assert:
- type: javascript
value: |
const wer = context.vars.metadata?.wer?.wer || 1;
wer < 0.1 // Less than 10% error
Speaker Count
tests:
- assert:
- type: javascript
value: |
const diarization = context.vars.metadata?.transcription?.diarization || [];
const uniqueSpeakers = new Set(diarization.map(s => s.speaker_id));
uniqueSpeakers.size === 2 // Expect 2 speakers
Language Support
ElevenLabs STT supports 30+ languages. Specify using ISO 639-1 codes:
config:
language: en # English
# language: es # Spanish
# language: fr # French
# language: de # German
# language: it # Italian
# language: pt # Portuguese
# language: ja # Japanese
# language: ko # Korean
# language: zh # Chinese
Auto-detection: Omit language to let the API detect the language automatically.
Cost Information
STT pricing is based on audio duration:
- Free tier: 1 hour/month
- Paid tiers:
$0.10 per minute ($0.00167 per second)
The provider automatically tracks and reports costs in the evaluation results.
Advanced Usage
Batch Transcription
prompts:
- audio/batch1.mp3
- audio/batch2.mp3
- audio/batch3.mp3
providers:
- id: elevenlabs:stt
config:
modelId: eleven_speech_to_text_v1
# Test all files with consistent assertions
tests:
- assert:
- type: cost
threshold: 0.10
- type: latency
threshold: 15000
Multi-language Testing
providers:
- id: elevenlabs:stt:english
config:
language: en
- id: elevenlabs:stt:spanish
config:
language: es
- id: elevenlabs:stt:autodetect
config:
# No language specified = auto-detect
prompts:
- audio/english_sample.mp3
- audio/spanish_sample.mp3
Accuracy Comparison
Compare transcription accuracy across different audio qualities:
prompts:
- audio/high_quality_48khz.wav
- audio/medium_quality_16khz.mp3
- audio/low_quality_8khz.mp3
providers:
- id: elevenlabs:stt
config:
calculateWER: true
referenceText: This is the expected transcription text
tests:
- description: High quality should have WER < 5%
vars:
audioFile: audio/high_quality_48khz.wav
assert:
- type: javascript
value: (context.vars.metadata?.wer?.wer || 1) < 0.05
- description: Medium quality should have WER < 10%
vars:
audioFile: audio/medium_quality_16khz.mp3
assert:
- type: javascript
value: (context.vars.metadata?.wer?.wer || 1) < 0.10
Troubleshooting
API Key Issues
# Verify your API key is set
echo $ELEVENLABS_API_KEY
# Or set it inline
ELEVENLABS_API_KEY=your_key promptfoo eval
Audio File Not Found
Error: Failed to read audio file: ENOENT: no such file or directory
Solution: Use absolute paths or paths relative to the config file:
prompts:
- /absolute/path/to/audio.mp3
- ./relative/path/to/audio.mp3
Unsupported Format
Error: Unsupported audio format
Solution: Convert your audio to a supported format (MP3, WAV, etc.) using tools like ffmpeg:
ffmpeg -i input.video -vn -acodec mp3 output.mp3
High WER on Clear Audio
If you're getting unexpectedly high WER:
- Check reference text - ensure it exactly matches the audio (including punctuation)
- Specify language - auto-detection may choose the wrong language
- Audio quality - ensure audio is clear with minimal background noise
- Normalization - WER calculation normalizes text (lowercase, removes punctuation)
API Reference
Config Options
| Option | Type | Default | Description |
|---|---|---|---|
modelId |
string | eleven_speech_to_text_v1 |
STT model to use |
language |
string | auto-detect | ISO 639-1 language code |
diarization |
boolean | false |
Enable speaker identification |
maxSpeakers |
number | - | Expected number of speakers |
audioFile |
string | - | Path to audio file |
audioFormat |
string | auto-detect | Audio format override |
referenceText |
string | - | Expected transcription for WER |
calculateWER |
boolean | false |
Calculate Word Error Rate |
baseUrl |
string | https://api.elevenlabs.io/v1 |
API endpoint |
timeout |
number | 120000 |
Request timeout (ms) |
retries |
number | 3 |
Number of retry attempts |
Related Examples
- ElevenLabs TTS - Text-to-Speech synthesis
- ElevenLabs Isolation - Audio cleanup quality comparison
Resources
FAQ
Common questions
Discussion
Questions & comments · 0
Sign In Sign in to leave a comment.