RIME RESEARCH

A new standard for conversational TTS benchmarks

Leaderboards focused on latency or "humanness" can't tell you how well a voice provider does the job it's tasked to do.

So we created three separate benchmarks: AlphaBench, SupportBench, and ContentBench. Each rank how well a vendor does at performing specific tasks.

AlphaBench

Did it read the tracking number right? We tested 580 clips of scripts of codes, IDs, and spellings past listeners.

SupportBench

Which voice would you rather hear on a support call? We ran 750 paired audio responses to listeners.

ContentBench

Which voice would you keep listening to? We ran 400 narration passages to listeners.

for support

SupportBench: the voice listeners chose for support calls

This benchmark judges voices reading order numbers, dates, amounts, addresses, and the sentences around them. Listeners heard the same script from two voices and picked the one they'd rather hear on a support call.

prefer Rime
about the same
prefer competitor
vs OpenAI
85%
6%
9%
vs Deepgram
62%
12%
26%
vs ElevenLabs
52%
8%
39%
vs Cartesia
51%
14%
35%

Listeners heard the same script from Rime and one competitor and picked the recording they'd rather hear on a support call. Each bar shows the share of the 750 scripts where they preferred Rime, heard about the same, or preferred the competitor.

for codes and IDs

AlphaBench: did it read the tracking number right?

580 scripts where one wrong character sends the package to the wrong house: confirmation codes, tracking numbers, flight numbers, name spellings. Listeners checked each clip against the script and flagged anything misread. Fewer flags is better.

ElevenLabs
6 clips
Cartesia
9 clips
Rime
10 clips
OpenAI
16 clips
Deepgram
84 clips

Out of 580 clips per voice, the number a majority of listeners flagged as misread. Fewer is better. Flags point reviewers at clips worth hearing; they aren't error rates.

omission
deepgram · confirmation code

“Your confirmation code is FFF888QXKF; you'll need it at check-in.”

3/3 listeners flagged · “forgot to pronounce the third F”

clip placeholder
early stop
deepgram · name spelling

“Please search for the customer under the spelling W-H-I-T-T-A-K-E-R.”

3/3 listeners flagged · “missing A-K-E-R”

clip placeholder
ambiguous
cartesia · cancellation code

“The cancellation code CKVZSO was sent to your email as well.”

3/3 listeners flagged · “the S sounded very low and off”

clip placeholder
for narration

ContentBench: narration is a different job

We had listeners review 400 passages of fiction and articles, with the same question: which voice would you rather listen to? The takeaway: The voices that win narration lose support calls, and that tradeoff is the reason this page has three charts instead of one score.

prefer Rime
about the same
prefer competitor
vs OpenAI
61%
7%
32%
vs Cartesia
58%
10%
32%
vs Deepgram
37%
8%
55%
vs ElevenLabs
37%
8%
56%

Same test with 400 narration passages. Each bar shows the share where listeners preferred Rime, heard about the same, or preferred the competitor. ElevenLabs and Deepgram come out ahead here; Cartesia and OpenAI don't.

get the data

Run it on your own voices

The scripts, configs, analysis, and raw judgments are all in the repo. Reproduce our numbers, or swap in your own voices and find out what wins on your workload.

Start the conversation