Speech To Text Tools

Independent of all 6 vendors11 services pricedvendor list prices, read 13 Sep 2026

One hour of your audio costs $0.15 at the cheapest of these 11 services and $1.44 at the dearest. That is the only number any of them will give you.

$0.15 AssemblyAI universal$1.44 Google v1_default

At 500 hours of audio a month, the distance between those two ends is $7,740 a year.

Which of them can actually hear your callers — the accents, the cross-talk, the fifty product names your business runs on — is on nobody's pricing page, and no vendor benchmark will settle it, because it was not run on your audio. Your own audio settles it. Send one real file through all 11 and most of this list eliminates itself in a few minutes; what comes back is the two or three worth your time, and the evidence for why.

What do you run in production today?at hours of audio a month
Run your audio across all 11One file, all 11 of them, and you watch them land before we ask who you are. The full report, and keeping it, need an account.

Input Source

Click or drag audio file here

Supports MP3, M4A, WAV, OGG

Up to 30 minutes of audio; longer recordings accepted up to 4 hours.

Got a reference transcript? Add it for word error rate.

A transcript you already know to be correct. Every provider's output gets scored against it — insertions, deletions and substitutions. Without one the run still happens; the report just says it cannot speak to accuracy.

Click to upload reference transcript

TXT or MD files

Configuration

5 models selected·
  • AssemblyAIus
  • AWSus-east-1
  • Azureeastus
  • Deepgramglobal
  • Googleus-central1
Advanced Options
Not supported by all selected

Provides a hint for the minimum and maximum number of expected speakers to improve diarization accuracy.

Boosts the recognition probability of specific words or phrases, such as proper nouns or domain-specific terms. Provide one phrase per line.

Note: Files and transcripts are not stored on our servers and are used only to complete your request. More features are coming.

Whether it hears your words

Accent, codec, cross-talk, and the fifty product names your business runs on. A vendor's benchmark was not run on your call centre at 8 kHz.

Measured on your audio. Needs a reference transcript for word error rate — most people have none on day one, so we start from where the services disagree with each other.

Whether it keeps up, and stays up

Throughput as a real-time factor rather than a millisecond count, end to end including the queue, and failures named rather than retried away.

Measured on your audio, and separately by our own probes against every service, around the clock, whether or not anyone uploads anything.

Whether it is still the model you chose

Services get replaced under a name that does not change. A choice made two years ago was made against a feature set that no longer exists.

We record what the vendor says it actually ran, for the vendors that say. Where they say nothing, the report says nothing.

ServiceAs publishedPer audio hourPer year at 500 h/moWhat would change this figure
1AssemblyAI universalus · batch$0.15 / hour$0.1500$900Universal-2, which is what our captured traffic shows the account is served. Universal-3.5 Pro is $0.21. Stereo is billed per channel.assemblyai.com/pricing
2Azure batch-defaulteastus · batch$0.18 / hour$0.1800$1,080Batch transcription, pay-as-you-go. Commitment tiers are cheaper.azure.microsoft.com/pricing/details/speech
3Deepgram nova_3global · batch$0.0043 / minute$0.2580$1,548Monolingual, pay-as-you-go. Multilingual is $0.0052 — 21% more — and true per-second billing, no minimum.deepgram.com/pricing
4Deepgram nova_2global · batch$0.0043 / minute$0.2580$1,548The weakest source on this page: the current pricing page no longer lists Nova-2, so this is the launch post plus a FAQ line promising unchanged rates.deepgram.com/learn/nova-2-speech-to-text-api
5Yandex stt3ru-central1 · async$0.0012418 / 15 s$0.2980$1,788Excludes VAT; the rouble list includes it. Minimum 15 seconds, and the unit is two channels — a 3-channel file costs double.aistudio.yandex.ru/docs/ru/speechkit/pricing
6Azure fast-defaulteastus · fast$0.36 / hour$0.3600$2,160Fast transcription: the same audio as batch, priced double for turnaround.azure.microsoft.com/pricing/details/speech
7AWS transcribeus-east-1 · batch$0.0001 / second$0.3600$2,160Identical in all five regions we deploy. Two-channel audio is not double-billed, unlike Google, AssemblyAI and Yandex.AWS price list API, published 2026-09-11
8Google chirp_3us · batch$0.016 / minute$0.9600$5,760First band. Volume discounts start at 500k minutes a month, cumulative per account. Each audio channel is billed separately.cloud.google.com/speech-to-text/pricing
9Google chirp_2us-central1 · batch$0.016 / minute$0.9600$5,760Same v2 standard-recognition meter as chirp_3, same volume bands.cloud.google.com/speech-to-text/pricing
10Azure short-defaulteastus · sync$1.00 / hour$1.0000$6,000Short audio, synchronous. Twenty times the batch fare for the same recogniser family.azure.microsoft.com/pricing/details/speech
11Google v1_defaultus · batch$0.024 / minute$1.4400$8,640The v1 meter, without data logging. Opting into data logging is $0.016; the first 60 minutes a month are free.cloud.google.com/speech-to-text/pricing

General-purpose batch and async transcription only. Streaming is a separate fare at every vendor and the specialty models — medical, call analytics — are a different product; putting them in this column would widen the spread by arithmetic rather than by fact. One service we support, AWS HealthScribe streaming, has no published fare in any region file, so it is absent rather than estimated. Pay-as-you-go throughout: committed-use and volume bands are cheaper, and they are the part of a vendor negotiation this page cannot see.

1

Give it one real file

An interview, a support call, a meeting. Long-form is the normal case here, not a ten-second clip. Nothing to fill in first — the fan-out starts as soon as the file lands, and you watch it happen.

2

It fans out to all 11 at once

Every service gets the same audio and the same options. They finish at wildly different times and some fail — both of those are results, and you watch them land one by one.

3

The report names your incumbent

Everything is measured against what you run today: cost delta first, then throughput, reliability, speakers, and accuracy where there is reference text to score it against.

4

You forward it to whoever asked

One page, one address, every figure sourced. This is the part behind the account, and the reason is the address: a report has to belong to someone to still be there when the question comes back next quarter. If it cannot survive being pasted into a Slack channel by someone who was not in the meeting, it did not answer the question.

Just need the text out of some videos?

Upload the pile, pick one model, see the price before anything starts. That door is over here. Same account, same balance, and if you later want to know whether you picked the right model, one file across all 11 is the way to find out.

Every price above is a vendor's own published figure, read from their pricing page on 13 Sep 2026 and converted to an audio hour here. None of it has been measured by us, which is the distinction the report exists to close: the sheet is what everyone can already read, and the run is what nobody else will do on your audio.

🚀