Hindi speech-to-text · published 5 August 2026
We didn't want a benchmark we won. We wanted one anyone can check. So we benchmarked Bodhi, Navana's Hindi speech-to-text, against five leading systems — Ringg, ElevenLabs, Deepgram, Sarvam and 60dB — on about 10,000 real Hindi clips, using the public benchmark data Ringg AI ↗ and 60dB ↗ published themselves, with the full method written out on this page, step by step.
We ran two tests. First the standard one — count the wrong words — which Bodhi wins. Then the one that matters if you're building voice agents: did the transcript keep the meaning? A wrong name, a wrong number or a lost "not" is what sends a live conversation off the rails — and on that test, no system beats Bodhi.
Our compiled Hindi set — audio, references and every system's transcript — is open on Hugging Face: Navana-AI/hindi-asr-benchmark ↗
Start here · listen to one clip
The speaker says: "The corporation has approved spoiling nature for money." Press play, then read what all six systems wrote down — and imagine a voice agent acting on each transcript.
Five systems, one identical word-count score — and not one of them kept who gave the approval. By the usual metric they're all equally good; anyone listening can tell they're not. That gap between same score and different outcome is what this whole benchmark is built to catch.
Test 1 · counting mistakes
The standard measure is Word Error Rate — the share of words a system gets wrong, so lower is better (12% means about one word in eight differs from what was said). On the neutral public rules, Bodhi is #1 at 11.8% — ahead of ElevenLabs (12.7%), Sarvam (15.5%), Ringg (17.5%), 60dB (21.2%) and Deepgram (21.6%).
But there's a catch worth understanding before you trust any vendor's headline number — including ours. The catch is what happens before the counting: every benchmark first "cleans up" the text — how it treats numbers, English words, spelling variants — and changing those rules reshuffles the ranking without any system getting one word more right. Take one vendor's transcripts from this very benchmark: under the neutral public rules they score 17.5%; under that vendor's own preferred rules, 7.4%. Same audio, same transcripts — a ten-point swing that comes entirely from who wrote the grading.
Here is exactly how that happens. Take one short sentence a system transcribed perfectly — it simply wrote the number and the English word in its own style:
| How the benchmark "cleans up" first | Still counts as wrong | WER |
|---|---|---|
| Count every difference literally | 50 ≠ पचास · online ≠ ऑनलाइन | 40% |
| …also treat 50 and पचास as the same number | online ≠ ऑनलाइन | 20% |
| …also treat English in either script as the same word | — nothing — | 0% |
Illustrative example. Same audio, same transcript — only the clean-up rules changed, and the score swung from 40% to 0%.
Here's the part that matters: none of these differences change the conversation. Whether the transcript reads "50" or "पचास", "online" or "ऑनलाइन", a voice agent acting on it does exactly the same thing — the speaker said the same number and the same word either way. These are spelling styles, not mistakes. A benchmark that counts them as errors is grading handwriting; one that forgives them selectively can quietly move the ranking. Meanwhile the errors that do change a conversation — a wrong name, a wrong amount — cost exactly one word each, same as a harmless spelling variant. Word-count treats them identically.
Every real benchmark makes a stack of these clean-up choices — numbers, English script, spelling marks, word spacing, even whether to drop a system's hardest clips. Each one quietly moves the ranking. So the question isn't "what's the WER" — it's "under whose rules."
So we didn't grade ourselves. We first rebuilt 60dB's and Ringg's own published scoreboards from the public data and matched them to the decimal — proving the scorer is honest — then scored every system, including Bodhi, under all three rule-books. Here's what happens to one and the same Bodhi as the rules change:
| Rule-book used for scoring | Bodhi's error rate | Where Bodhi finishes |
|---|---|---|
| Neutral public rules (strictest) | 11.8% | #1 of 6 — ahead of everyone |
| Ringg's own rules (most forgiving) | 8.1% | #2 — 0.6 behind Ringg, ahead of the rest |
| Fairest rules, identical for everyone | 7.6% | #1 on the per-set average, by a whisker |
Same model, same transcripts — three different numbers. A rule-book moves every system's score up or down together, so never compare scores across rule-books — compare the finishing order inside one. Bodhi finishes at the top of all three.
Under the fairest rules — the same clean-up applied to everyone — Bodhi comes out on top, edging Ringg 7.57% to 7.58% on the per-set average, with everyone else more than a full point behind:
But look at the size of that lead over Ringg: 0.01. Two systems this close can't be separated by counting words — and as the example above showed, word-count can't tell a harmless spelling variant from a wrong name that derails a call. To find out which system you can actually trust in a live conversation, you have to stop counting words and start checking meaning.
Test 2 · understanding
For a voice agent, most transcript differences are harmless — but a wrong name greets the wrong customer, a wrong number confirms the wrong amount, and a lost "not" flips an instruction into its opposite. Those are the errors that end conversations. So we had an AI judge — Google's Gemini 2.5 Pro — read every transcript and rate whether it kept what was actually said, ignoring harmless spelling differences. It never saw brand names — each system was anonymized and shuffled on every clip — so no vendor's style could sway it.
Judged this way, no system beats Bodhi. Against most of the field it isn't close: clip for clip, the judge rated Bodhi better than Sarvam 312 times vs 109, better than Deepgram 312 vs 75, and better than 60dB 377 vs 53. At the very top, Bodhi is level with the strongest rival, ElevenLabs (a statistical coin-flip, p=0.48) — and it breaks the tie that word-counting couldn't: against Ringg, Bodhi is rated better 178 times vs 110, an edge the statistics say is real, not luck (p<0.0001).
Bodhi vs each rival on the same clips, rated by the same blind judge · numbers are clip counts. Full meaning-score breakdown below.
Listen for yourself
Browse real clips from the benchmark — press play, then read what every system wrote. The clearest cases are clips where Bodhi and several rivals earn the exact same word-count score, yet only Bodhi keeps the meaning — a wrong name, a flipped word: exactly the slip that would send an agent down the wrong path mid-call.
Trust · nothing hidden
Every number here is re-derivable from public data — we even reproduced Ringg's and 60dB's own published tables before adding Bodhi. Here's the full picture: the different grading rules and why our result holds under all of them, how we made sure the AI judge was fair, and the complete method — from downloading the raw data to the final scores.
Because each set of clean-up rules can shift the ranking, we scored every system three ways. The takeaway: Bodhi stays near the top no matter which rules you use — its ranking doesn't depend on who wrote the test.
The three rule-sets in full (WER %, lower is better; ● = best in row):
60dB way = neutral public rules (Bodhi #1). Ringg way = Ringg's own rules (Bodhi #2, close behind Ringg). Navana way = a fair best-of-both applied to everyone (Bodhi and Ringg tie for #1; Bodhi wins the per-set average).
The judge (Google's Gemini 2.5 Pro) was anchored to the human-written reference, so it couldn't invent its own answer, and it ran deterministically (temperature 0). Two checks confirm the scores are a measurement, not a hunch: re-running with a fresh anonymized shuffle agrees 97.1% of the time, and there's no bias toward any position on the list. For the record, on the overall score the top three (ElevenLabs, Ringg, Bodhi) overlap within their margins of error — a genuine tie — with Bodhi vs ElevenLabs a coin-flip (p=0.48) and Bodhi's head-to-head edge over Ringg at p<0.0001.
System instruction (sent to gemini-2.5-pro, temperature 0)
You are a meticulous Hindi (Devanagari) linguist auditing speech-recognition output. You are given a GROUND-TRUTH reference sentence and several anonymous candidate transcripts of the SAME audio. Rate how faithfully EACH candidate preserves the MEANING of the reference. Judge MEANING, not spelling. IGNORE these — they are NOT errors: - nukta / matra spelling variants (ज़ vs ज, क़ vs क), chandrabindu vs anusvara (ँ/ं) - word spacing / compound splits (हो कर vs होकर) - a number written as digits vs spelled out (25 vs पच्चीस) IF the value is the same - an English word written in Latin vs Devanagari (police vs पुलिस) IF it is the same word - punctuation, case, extra/missing whitespace PENALISE only genuine recognition errors that change or lose meaning: - a content word wrong, dropped, or invented (hallucinated) - a named entity wrong (person/place/org) - a NUMBER with a different value - a negation or key modifier lost/flipped Severity for each candidate: 0 PERFECT — full meaning; any differences are purely cosmetic (see IGNORE list) 1 COSMETIC — only cosmetic differences; meaning fully intact (treat 0 and 1 alike) 2 MINOR — a trivial slip (a filler/particle); the gist is fully recoverable 3 MAJOR — a content word / entity / number wrong or dropped; meaning changed 4 SEVERE — meaning largely wrong or missing error_type: one of none, cosmetic, minor, content_word, named_entity, number, omission, hallucination, negation, severe. Return exactly one rating object per candidate label provided. Keep note to <= 12 words.
1 · Getting the data. Ringg AI and 60dB each published a Hindi benchmark dataset on Hugging Face (RinggAI/ASR-Benchmarking-Dataset ↗ and SkunkWorkLabs/hindi-asr-benchmark ↗). They turn out to be the same corpus — 60dB republished Ringg's references and provider transcripts byte-for-byte, adding its own model's column. It holds roughly 10,000 real Hindi clips from six public test sets (Common Voice, FLEURS, IndicTTS, Kathbath, Kathbath-noisy and MUCS), each clip carrying a human-written reference transcript and the transcripts produced by Sarvam, Deepgram, ElevenLabs, Ringg and 60dB. We downloaded both datasets, verified they match, and kept the 9,929 clips present in both. We never re-ran any competitor's system ourselves — every rival is represented by transcripts these vendors published — so no rival's result depends on how well we operated their product.
2 · Adding Bodhi. We ran the same audio through Bodhi once and froze the output — no retries, no manual fixes, no dropping of inconvenient clips.
3 · Rebuilding both vendors' graders — and proving them correct. Before words can be counted, every benchmark first "cleans up" the text. The two vendors' recipes differ a lot:
We re-implemented both recipes from public artifacts, then proved the re-implementations by reproducing each vendor's own published score tables from raw data: 60dB's 24-cell table to within 0.05 of a WER point on average (21 of 24 cells exact to the decimal), and Ringg's two tables to within 0.07 and 0.27, with every residual gap traced to a documented cause. Only after our grader could reproduce their numbers did we trust it to grade anyone.
4 · Building the fair rule-book. We then combined the defensible parts of both recipes into one, applied identically to the reference and to every system on every clip:
Whatever survives this clean-up is a genuine recognition error, not a style difference.
5 · Counting. The industry-standard word-error calculation — wrong plus missing plus invented words, divided by the number of words actually said — reported two ways: all clips pooled together, and the average of the six test sets. That way no aggregation choice can flatter anyone. These are the tables above.
6 · Checking meaning. On a random sample of roughly 700 clips (every system judged on the same clips), the AI judge — Google's Gemini 2.5 Pro, run deterministically at temperature 0 — was shown the human reference plus all six transcripts, with vendor names removed and the order reshuffled on every clip so nothing could be favored. It rated each transcript on a 0–4 scale (0–1 = meaning fully kept, 2 = minor slip, 3–4 = meaning changed) and named the error type. "Kept the meaning" on this page means a rating of 0 or 1. Head-to-head results are decided clip by clip; ties are shown but sit out of the statistics; the p-values come from a standard sign test; and re-running the whole judgment with a fresh shuffle agreed with itself 97.1% of the time. The exact prompt is in the drawer above.
Beyond Hindi · ten languages
This report covers Hindi, but Bodhi is built the same accuracy-first way across ten languages — so a voice agent that stays on track in Hindi stays on track for the rest of your callers too.
Want this accuracy on your own calls? Bring us a sample of your audio and we'll show you what Bodhi hears — book a meeting with the team.
Book a meeting with us ↗