Your reference transcriptEvery provider scored against itstrict and normalised, side by side
Word error rate is the only accuracy number worth arguing about, and it is also the easiest one to quietly flatter. Score “twenty three” against “23” and you have either one substitution or none, depending entirely on how you normalise before counting — same audio, same model, two different WERs and no dishonesty required. A vendor quotes whichever of those numbers reads best. We give you all of them at once, on your transcript, for every provider in the same run.
Your audio, your transcript, every provider
Click or drag audio file here
Supports MP3, M4A, WAV, OGG
Up to 30 minutes of audio; longer recordings accepted up to 4 hours.
A transcript you already know to be correct. Every provider's output gets scored against it — insertions, deletions and substitutions. Without one the run still happens; the report just says it cannot speak to accuracy.
Click to upload reference transcript
TXT or MD files
Provides a hint for the minimum and maximum number of expected speakers to improve diarization accuracy.
Boosts the recognition probability of specific words or phrases, such as proper nouns or domain-specific terms. Provide one phrase per line.
Note: Files and transcripts are not stored on our servers and are used only to complete your request. More features are coming.
What one WER number will not tell you
Which errors, not how many
A 12% WER made of dropped filler words and a 12% WER made of mangled product names are the same number and completely different problems. One you ship; the other loses a customer.
Insertions, deletions and substitutions are broken out per provider, against your reference, so you can look at what was actually got wrong.
How much the normalisation decided
Case, punctuation, numbers written as words, currency, dates. Strip more and every model looks better; strip less and everything looks broken. The choice moves the score more than the models differ from each other.
Strict and normalised are both returned, per provider, so you can see how much of any lead is real and how much is the scoring.
Whether one file settles it
One recording is one accent, one codec, one room. It will tell you which providers are unusable on your audio, which is most of the value — it will not tell you the ranking of the two that survive.
For that, run a handful of files that differ in the way your real traffic differs. The score has to earn the confidence you put on it.
What happens
Give it the audio and the transcript
The transcript is one you already know to be correct — a human pass, a published caption file, a script that was read. Plain text or markdown, pasted or uploaded. It is the ground truth everything else is measured against, so it is worth being fussy about.
Every provider transcribes the same file
The same audio and the same options to all of them at once. They finish at different times and some fail — both are results, and you watch them land one by one.
Each output is scored against your text
Word error rate and character error rate, strict and normalised, per provider, with the errors broken out rather than summed into a single percentage.
You compare the transcripts, not just the scores
The number tells you where to look; the diff tells you whether it matters. A model that loses punctuation and a model that invents a speaker score similarly and are not similar.
Normalisation is Nvidia NeMo's inverse text normalisation, which is why the normalised figure is comparable across languages rather than being our own regex — how that works and what it changes. Cost and speed for every provider are on the main comparison, which does the same fan-out without needing a reference transcript.