Speech To Text Tools

Your reference transcriptEvery provider scored against itstrict and normalised, side by side

You already know what the audio says. Find out which of them can hear it.

Word error rate is the only accuracy number worth arguing about, and it is also the easiest one to quietly flatter. Score “twenty three” against “23” and you have either one substitution or none, depending entirely on how you normalise before counting — same audio, same model, two different WERs and no dishonesty required. A vendor quotes whichever of those numbers reads best. We give you all of them at once, on your transcript, for every provider in the same run.

Score them against your transcriptNo reference transcript? Start hereAudio plus a transcript you trust. The full report, and keeping it, need an account.

Input Source

Click or drag audio file here

Supports MP3, M4A, WAV, OGG

Up to 30 minutes of audio; longer recordings accepted up to 4 hours.

Got a reference transcript? Add it for word error rate.

A transcript you already know to be correct. Every provider's output gets scored against it — insertions, deletions and substitutions. Without one the run still happens; the report just says it cannot speak to accuracy.

Click to upload reference transcript

TXT or MD files

Configuration

5 models selected·
  • AssemblyAIus
  • AWSus-east-1
  • Azureeastus
  • Deepgramglobal
  • Googleus-central1
Advanced Options
Not supported by all selected

Provides a hint for the minimum and maximum number of expected speakers to improve diarization accuracy.

Boosts the recognition probability of specific words or phrases, such as proper nouns or domain-specific terms. Provide one phrase per line.

Note: Files and transcripts are not stored on our servers and are used only to complete your request. More features are coming.

Which errors, not how many

A 12% WER made of dropped filler words and a 12% WER made of mangled product names are the same number and completely different problems. One you ship; the other loses a customer.

Insertions, deletions and substitutions are broken out per provider, against your reference, so you can look at what was actually got wrong.

How much the normalisation decided

Case, punctuation, numbers written as words, currency, dates. Strip more and every model looks better; strip less and everything looks broken. The choice moves the score more than the models differ from each other.

Strict and normalised are both returned, per provider, so you can see how much of any lead is real and how much is the scoring.

Whether one file settles it

One recording is one accent, one codec, one room. It will tell you which providers are unusable on your audio, which is most of the value — it will not tell you the ranking of the two that survive.

For that, run a handful of files that differ in the way your real traffic differs. The score has to earn the confidence you put on it.

1

Give it the audio and the transcript

The transcript is one you already know to be correct — a human pass, a published caption file, a script that was read. Plain text or markdown, pasted or uploaded. It is the ground truth everything else is measured against, so it is worth being fussy about.

2

Every provider transcribes the same file

The same audio and the same options to all of them at once. They finish at different times and some fail — both are results, and you watch them land one by one.

3

Each output is scored against your text

Word error rate and character error rate, strict and normalised, per provider, with the errors broken out rather than summed into a single percentage.

4

You compare the transcripts, not just the scores

The number tells you where to look; the diff tells you whether it matters. A model that loses punctuation and a model that invents a speaker score similarly and are not similar.

Normalisation is Nvidia NeMo's inverse text normalisation, which is why the normalised figure is comparable across languages rather than being our own regex — how that works and what it changes. Cost and speed for every provider are on the main comparison, which does the same fan-out without needing a reference transcript.

🚀