Pranix Aaria
What we measured, and how
People in Telangana mix Telugu, Hindi and English in one sentence, so we measure each language separately, on recordings of real people, and publish only what we measured. The figures are word error rates: the share of words a recogniser got wrong. Lower is better.
1 · Cloud speech recognition
Measured 30 September – 1 October 2026
| Recogniser | Telugu | Hindi | English |
|---|---|---|---|
| Sarvam v4 with script mode and a 50-word boost list | 13.0% | 6.2% | not measured |
| Sarvam v3 | 16.4% | 8.4% | not measured |
| AI4Bharat IndicConformer (open source, MIT, run on a CPU) | 16.4% | 10.6% | not offered |
| NVIDIA Canary | not offered | 10.7% | 4.8% |
| Bhashini | 18.1% | not published* | not measured |
- Every recogniser was scored on the same recordings of real people reading sentences (150 in this comparison; 49 Telugu and 50 Hindi in the IndicConformer run) with the same scorer.
- “Not offered” means the service has no model for that language. We do not score it; a 0% would be false.
- The Sarvam v4 result combines three changes (v4, script mode and the boost list), which we have not yet measured separately.
- *Bhashini's Hindi result was recorded only as level with the others, not as a number, so it is left out.
- Cloud round-trip time: half of requests within 4.4 s, 95% within 7.6 s, on a free-tier cloud server with no GPU.
The two result sets on this page come from different tests and cannot be compared with each other. Each number carries its own test.
2 · Aaria Edge: on the phone, no internet
Measured 30 September 2026
| Language | Model on the phone | Word error | Seconds per 3-s clip (PC) |
|---|---|---|---|
| English | Moonshine tiny, int8 (MIT) | 14.3% | 0.17 |
| Hindi | IndicConformer, int8 (MIT) | 12.3% | 0.40 |
| Telugu | IndicConformer, int8 (MIT) | 23.0% | 0.80 |
- Test set: 30 read sentences per language from Google's public FLEURS test set, spoken by real people. That is long read speech, used to choose a model; on the phone Aaria only needs to recognise short commands.
- Speeds were measured on a PC. Phone speeds will be published after tests on a low-cost and a mid-range phone, not before.
- Telugu is the weakest result on both tests, and improving it is our main focus.
Rules behind every number
- Only numbers measured on the build described are published.
- A number we do not have is shown as “not measured”, never as zero.
- Computer-generated speech is never scored as if it were a person speaking.
- Languages not yet checked by a native speaker have no published figure.