Pranix Aaria

What we measured, and how

People in Telangana mix Telugu, Hindi and English in one sentence, so we measure each language separately, on recordings of real people, and publish only what we measured. The figures are word error rates: the share of words a recogniser got wrong. Lower is better.

1 · Cloud speech recognition

Measured 30 September – 1 October 2026

RecogniserTeluguHindiEnglish
Sarvam v4 with script mode and a 50-word boost list13.0%6.2%not measured
Sarvam v316.4%8.4%not measured
AI4Bharat IndicConformer (open source, MIT, run on a CPU)16.4%10.6%not offered
NVIDIA Canarynot offered10.7%4.8%
Bhashini18.1%not published*not measured
  • Every recogniser was scored on the same recordings of real people reading sentences (150 in this comparison; 49 Telugu and 50 Hindi in the IndicConformer run) with the same scorer.
  • “Not offered” means the service has no model for that language. We do not score it; a 0% would be false.
  • The Sarvam v4 result combines three changes (v4, script mode and the boost list), which we have not yet measured separately.
  • *Bhashini's Hindi result was recorded only as level with the others, not as a number, so it is left out.
  • Cloud round-trip time: half of requests within 4.4 s, 95% within 7.6 s, on a free-tier cloud server with no GPU.

The two result sets on this page come from different tests and cannot be compared with each other. Each number carries its own test.

2 · Aaria Edge: on the phone, no internet

Measured 30 September 2026

LanguageModel on the phoneWord errorSeconds per 3-s clip (PC)
EnglishMoonshine tiny, int8 (MIT)14.3%0.17
HindiIndicConformer, int8 (MIT)12.3%0.40
TeluguIndicConformer, int8 (MIT)23.0%0.80
  • Test set: 30 read sentences per language from Google's public FLEURS test set, spoken by real people. That is long read speech, used to choose a model; on the phone Aaria only needs to recognise short commands.
  • Speeds were measured on a PC. Phone speeds will be published after tests on a low-cost and a mid-range phone, not before.
  • Telugu is the weakest result on both tests, and improving it is our main focus.

Rules behind every number

  • Only numbers measured on the build described are published.
  • A number we do not have is shown as “not measured”, never as zero.
  • Computer-generated speech is never scored as if it were a person speaking.
  • Languages not yet checked by a native speaker have no published figure.
Ask me — I'm Aaria, your assistant ✨