A new standard for conversational TTS benchmarks
Leaderboards focused on latency or "humanness" can't tell you how well a voice provider does the job it's tasked to do.
So we created three separate benchmarks: AlphaBench, SupportBench, and ContentBench. Each rank how well a vendor does at performing specific tasks.
AlphaBench
Did it read the tracking number right? We tested 580 clips of scripts of codes, IDs, and spellings past listeners.
SupportBench
Which voice would you rather hear on a support call? We ran 750 paired audio responses to listeners.
ContentBench
Which voice would you keep listening to? We ran 400 narration passages to listeners.
SupportBench: the voice listeners chose for support calls
This benchmark judges voices reading order numbers, dates, amounts, addresses, and the sentences around them. Listeners heard the same script from two voices and picked the one they'd rather hear on a support call.
Listeners heard the same script from Rime and one competitor and picked the recording they'd rather hear on a support call. Each bar shows the share of the 750 scripts where they preferred Rime, heard about the same, or preferred the competitor.
AlphaBench: did it read the tracking number right?
580 scripts where one wrong character sends the package to the wrong house: confirmation codes, tracking numbers, flight numbers, name spellings. Listeners checked each clip against the script and flagged anything misread. Fewer flags is better.
Out of 580 clips per voice, the number a majority of listeners flagged as misread. Fewer is better. Flags point reviewers at clips worth hearing; they aren't error rates.
“Your confirmation code is FFF888QXKF; you'll need it at check-in.”
3/3 listeners flagged · “forgot to pronounce the third F”
“Please search for the customer under the spelling W-H-I-T-T-A-K-E-R.”
3/3 listeners flagged · “missing A-K-E-R”
“The cancellation code CKVZSO was sent to your email as well.”
3/3 listeners flagged · “the S sounded very low and off”
ContentBench: narration is a different job
We had listeners review 400 passages of fiction and articles, with the same question: which voice would you rather listen to? The takeaway: The voices that win narration lose support calls, and that tradeoff is the reason this page has three charts instead of one score.
Same test with 400 narration passages. Each bar shows the share where listeners preferred Rime, heard about the same, or preferred the competitor. ElevenLabs and Deepgram come out ahead here; Cartesia and OpenAI don't.
Run it on your own voices
The scripts, configs, analysis, and raw judgments are all in the repo. Reproduce our numbers, or swap in your own voices and find out what wins on your workload.