Navana Bodhiby Navana ← All posts Get in touch ↗

Hindi speech-to-text · published 5 August 2026

Whose voice AI really hears Hindi best? We built a benchmark anyone can check.

We didn't want a benchmark we won. We wanted one anyone can check. So we benchmarked Bodhi, Navana's Hindi speech-to-text, against five leading systems — Ringg, ElevenLabs, Deepgram, Sarvam and 60dB — on about 10,000 real Hindi clips, using the public benchmark data Ringg AI ↗ and 60dB ↗ published themselves, with the full method written out on this page, step by step.

We ran two tests. First the standard one — count the wrong words — which Bodhi wins. Then the one that matters if you're building voice agents: did the transcript keep the meaning? A wrong name, a wrong number or a lost "not" is what sends a live conversation off the rails — and on that test, no system beats Bodhi.

Our compiled Hindi set — audio, references and every system's transcript — is open on Hugging Face: Navana-AI/hindi-asr-benchmark ↗

The results, in short
Fewest wrong words. On the public scoring rules Bodhi is #1 — 11.8% error vs ElevenLabs 12.7%, Sarvam 15.5%, Ringg 17.5%, 60dB 21.2%, Deepgram 21.6%. Other rule-books forgive more, so every score looks better under them (Bodhi's becomes 7.6%) — but Bodhi stays at the top under all of them.
Meaning kept, head-to-head. A blind AI judge compared transcripts clip by clip: Bodhi beats Sarvam 312–109, Deepgram 312–75 and 60dB 377–53, edges Ringg 178–110 — and no system comes out ahead of it.
Everything is checkable. Public data, a written step-by-step method, clips you can listen to — every number on this page can be re-derived by anyone.

Start here · listen to one clip

One wrong word can change a whole conversation.

The speaker says: "The corporation has approved spoiling nature for money." Press play, then read what all six systems wrote down — and imagine a voice agent acting on each transcript.

Five systems, one identical word-count score — and not one of them kept who gave the approval. By the usual metric they're all equally good; anyone listening can tell they're not. That gap between same score and different outcome is what this whole benchmark is built to catch.

Test 1 · counting mistakes

First we counted mistakes. Bodhi made the fewest.

The standard measure is Word Error Rate — the share of words a system gets wrong, so lower is better (12% means about one word in eight differs from what was said). On the neutral public rules, Bodhi is #1 at 11.8% — ahead of ElevenLabs (12.7%), Sarvam (15.5%), Ringg (17.5%), 60dB (21.2%) and Deepgram (21.6%).

But there's a catch worth understanding before you trust any vendor's headline number — including ours. The catch is what happens before the counting: every benchmark first "cleans up" the text — how it treats numbers, English words, spelling variants — and changing those rules reshuffles the ranking without any system getting one word more right. Take one vendor's transcripts from this very benchmark: under the neutral public rules they score 17.5%; under that vendor's own preferred rules, 7.4%. Same audio, same transcripts — a ten-point swing that comes entirely from who wrote the grading.

Here is exactly how that happens. Take one short sentence a system transcribed perfectly — it simply wrote the number and the English word in its own style:

What was saidकरीब 50 लोग online आए
System wroteकरीब पचास लोग ऑनलाइन आए
How the benchmark "cleans up" firstStill counts as wrongWER
Count every difference literally50 ≠ पचास · online ≠ ऑनलाइन40%
…also treat 50 and पचास as the same numberonline ≠ ऑनलाइन20%
…also treat English in either script as the same word— nothing —0%

Illustrative example. Same audio, same transcript — only the clean-up rules changed, and the score swung from 40% to 0%.

Here's the part that matters: none of these differences change the conversation. Whether the transcript reads "50" or "पचास", "online" or "ऑनलाइन", a voice agent acting on it does exactly the same thing — the speaker said the same number and the same word either way. These are spelling styles, not mistakes. A benchmark that counts them as errors is grading handwriting; one that forgives them selectively can quietly move the ranking. Meanwhile the errors that do change a conversation — a wrong name, a wrong amount — cost exactly one word each, same as a harmless spelling variant. Word-count treats them identically.

Every real benchmark makes a stack of these clean-up choices — numbers, English script, spelling marks, word spacing, even whether to drop a system's hardest clips. Each one quietly moves the ranking. So the question isn't "what's the WER" — it's "under whose rules."

So we didn't grade ourselves. We first rebuilt 60dB's and Ringg's own published scoreboards from the public data and matched them to the decimal — proving the scorer is honest — then scored every system, including Bodhi, under all three rule-books. Here's what happens to one and the same Bodhi as the rules change:

Rule-book used for scoringBodhi's error rateWhere Bodhi finishes
Neutral public rules (strictest)11.8%#1 of 6 — ahead of everyone
Ringg's own rules (most forgiving)8.1%#2 — 0.6 behind Ringg, ahead of the rest
Fairest rules, identical for everyone7.6%#1 on the per-set average, by a whisker

Same model, same transcripts — three different numbers. A rule-book moves every system's score up or down together, so never compare scores across rule-books — compare the finishing order inside one. Bodhi finishes at the top of all three.

Under the fairest rules — the same clean-up applied to everyone — Bodhi comes out on top, edging Ringg 7.57% to 7.58% on the per-set average, with everyone else more than a full point behind:

Word Error Rate under the fairest rules · per-set average · Bodhi #1 · lower is better · ~10,000 Hindi clips

But look at the size of that lead over Ringg: 0.01. Two systems this close can't be separated by counting words — and as the example above showed, word-count can't tell a harmless spelling variant from a wrong name that derails a call. To find out which system you can actually trust in a live conversation, you have to stop counting words and start checking meaning.

Test 2 · understanding

Then we asked the question that matters: did it understand what was said?

For a voice agent, most transcript differences are harmless — but a wrong name greets the wrong customer, a wrong number confirms the wrong amount, and a lost "not" flips an instruction into its opposite. Those are the errors that end conversations. So we had an AI judge — Google's Gemini 2.5 Pro — read every transcript and rate whether it kept what was actually said, ignoring harmless spelling differences. It never saw brand names — each system was anonymized and shuffled on every clip — so no vendor's style could sway it.

Judged this way, no system beats Bodhi. Against most of the field it isn't close: clip for clip, the judge rated Bodhi better than Sarvam 312 times vs 109, better than Deepgram 312 vs 75, and better than 60dB 377 vs 53. At the very top, Bodhi is level with the strongest rival, ElevenLabs (a statistical coin-flip, p=0.48) — and it breaks the tie that word-counting couldn't: against Ringg, Bodhi is rated better 178 times vs 110, an edge the statistics say is real, not luck (p<0.0001).

clips where Bodhi kept the meaning better tie — both did equally well clips where the rival did better

Bodhi vs each rival on the same clips, rated by the same blind judge · numbers are clip counts. Full meaning-score breakdown below.

Listen for yourself

Play real clips and compare what every system heard.

Browse real clips from the benchmark — press play, then read what every system wrote. The clearest cases are clips where Bodhi and several rivals earn the exact same word-count score, yet only Bodhi keeps the meaning — a wrong name, a flipped word: exactly the slip that would send an agent down the wrong path mid-call.

got the meaning right minor slip changed the meaning
Loading clips…

Trust · nothing hidden

Don't take our word for it — check everything yourself.

Every number here is re-derivable from public data — we even reproduced Ringg's and 60dB's own published tables before adding Bodhi. Here's the full picture: the different grading rules and why our result holds under all of them, how we made sure the AI judge was fair, and the complete method — from downloading the raw data to the final scores.

The full scores — three different rule-books, Bodhi at the top of eachtables + chart

Because each set of clean-up rules can shift the ranking, we scored every system three ways. The takeaway: Bodhi stays near the top no matter which rules you use — its ranking doesn't depend on who wrote the test.

Each line is one system's error rate under three different rule-sets · Bodhi stays low across all of them · lower is better

The three rule-sets in full (WER %, lower is better; ● = best in row):

60dB way = neutral public rules (Bodhi #1). Ringg way = Ringg's own rules (Bodhi #2, close behind Ringg). Navana way = a fair best-of-both applied to everyone (Bodhi and Ringg tie for #1; Bodhi wins the per-set average).

How we made sure the AI judge was fairjudge details

The judge (Google's Gemini 2.5 Pro) was anchored to the human-written reference, so it couldn't invent its own answer, and it ran deterministically (temperature 0). Two checks confirm the scores are a measurement, not a hunch: re-running with a fresh anonymized shuffle agrees 97.1% of the time, and there's no bias toward any position on the list. For the record, on the overall score the top three (ElevenLabs, Ringg, Bodhi) overlap within their margins of error — a genuine tie — with Bodhi vs ElevenLabs a coin-flip (p=0.48) and Bodhi's head-to-head edge over Ringg at p<0.0001.

Overall "got the meaning right" rate — bar; bracket = 95% confidence range · higher is better:
See the exact prompt the judge was givenverbatim

System instruction (sent to gemini-2.5-pro, temperature 0)

You are a meticulous Hindi (Devanagari) linguist auditing speech-recognition output.
You are given a GROUND-TRUTH reference sentence and several anonymous candidate
transcripts of the SAME audio. Rate how faithfully EACH candidate preserves the
MEANING of the reference.

Judge MEANING, not spelling. IGNORE these — they are NOT errors:
  - nukta / matra spelling variants (ज़ vs ज, क़ vs क), chandrabindu vs anusvara (ँ/ं)
  - word spacing / compound splits (हो कर vs होकर)
  - a number written as digits vs spelled out (25 vs पच्चीस) IF the value is the same
  - an English word written in Latin vs Devanagari (police vs पुलिस) IF it is the same word
  - punctuation, case, extra/missing whitespace

PENALISE only genuine recognition errors that change or lose meaning:
  - a content word wrong, dropped, or invented (hallucinated)
  - a named entity wrong (person/place/org)
  - a NUMBER with a different value
  - a negation or key modifier lost/flipped

Severity for each candidate:
  0 PERFECT  — full meaning; any differences are purely cosmetic (see IGNORE list)
  1 COSMETIC — only cosmetic differences; meaning fully intact  (treat 0 and 1 alike)
  2 MINOR    — a trivial slip (a filler/particle); the gist is fully recoverable
  3 MAJOR    — a content word / entity / number wrong or dropped; meaning changed
  4 SEVERE   — meaning largely wrong or missing

error_type: one of none, cosmetic, minor, content_word, named_entity, number,
omission, hallucination, negation, severe.
Return exactly one rating object per candidate label provided. Keep note to <= 12 words.
The complete method — every step, in detailstep by step

1 · Getting the data. Ringg AI and 60dB each published a Hindi benchmark dataset on Hugging Face (RinggAI/ASR-Benchmarking-Dataset ↗ and SkunkWorkLabs/hindi-asr-benchmark ↗). They turn out to be the same corpus — 60dB republished Ringg's references and provider transcripts byte-for-byte, adding its own model's column. It holds roughly 10,000 real Hindi clips from six public test sets (Common Voice, FLEURS, IndicTTS, Kathbath, Kathbath-noisy and MUCS), each clip carrying a human-written reference transcript and the transcripts produced by Sarvam, Deepgram, ElevenLabs, Ringg and 60dB. We downloaded both datasets, verified they match, and kept the 9,929 clips present in both. We never re-ran any competitor's system ourselves — every rival is represented by transcripts these vendors published — so no rival's result depends on how well we operated their product.

2 · Adding Bodhi. We ran the same audio through Bodhi once and froze the output — no retries, no manual fixes, no dropping of inconvenient clips.

3 · Rebuilding both vendors' graders — and proving them correct. Before words can be counted, every benchmark first "cleans up" the text. The two vendors' recipes differ a lot:

  • 60dB's recipe — strip punctuation; expand digits into Hindi words (50 → पचास; long numbers digit-by-digit, 1976 → एक नौ सात छह); keep English words in English script; canonicalize the text encoding so two computer spellings of the same letter don't count as an error. Its cleaned-up reference text is published, so every step can be audited.
  • Ringg's recipe — keep numbers as digits; convert English words to Devanagari with a transliteration model (we regenerated the real conversion table for all 3,004 English tokens in the corpus); drop dot-marks (ज़ → ज) and fold nasal spelling variants (लम्बा = लंबा); forgive words written joined vs split (हो कर = होकर); and drop, per provider, the clips its transliterator can't process.

We re-implemented both recipes from public artifacts, then proved the re-implementations by reproducing each vendor's own published score tables from raw data: 60dB's 24-cell table to within 0.05 of a WER point on average (21 of 24 cells exact to the decimal), and Ringg's two tables to within 0.07 and 0.27, with every residual gap traced to a documented cause. Only after our grader could reproduce their numbers did we trust it to grade anyone.

4 · Building the fair rule-book. We then combined the defensible parts of both recipes into one, applied identically to the reference and to every system on every clip:

  • numbers matched by value — 25 and पच्चीस are the same number;
  • English accepted in either script — online and ऑनलाइन are the same word, via the same published conversion table;
  • spelling-mark variants folded — ज़ = ज, ँ = ं, लम्बा = लंबा;
  • joined-vs-split words forgiven only when the letters are identical (हो कर = होकर) — an actually-wrong word can never be forgiven this way;
  • all 9,929 clips scored for every system — no clips dropped for anyone.

Whatever survives this clean-up is a genuine recognition error, not a style difference.

5 · Counting. The industry-standard word-error calculation — wrong plus missing plus invented words, divided by the number of words actually said — reported two ways: all clips pooled together, and the average of the six test sets. That way no aggregation choice can flatter anyone. These are the tables above.

6 · Checking meaning. On a random sample of roughly 700 clips (every system judged on the same clips), the AI judge — Google's Gemini 2.5 Pro, run deterministically at temperature 0 — was shown the human reference plus all six transcripts, with vendor names removed and the order reshuffled on every clip so nothing could be favored. It rated each transcript on a 0–4 scale (0–1 = meaning fully kept, 2 = minor slip, 3–4 = meaning changed) and named the error type. "Kept the meaning" on this page means a rating of 0 or 1. Head-to-head results are decided clip by clip; ties are shown but sit out of the statistics; the p-values come from a standard sign test; and re-running the whole judgment with a fresh shuffle agreed with itself 97.1% of the time. The exact prompt is in the drawer above.

Beyond Hindi · ten languages

And it's not just Hindi — Bodhi works in ten languages.

This report covers Hindi, but Bodhi is built the same accuracy-first way across ten languages — so a voice agent that stays on track in Hindi stays on track for the rest of your callers too.

Want this accuracy on your own calls? Bring us a sample of your audio and we'll show you what Bodhi hears — book a meeting with the team.

Book a meeting with us ↗