Hugging Face and Voice Arena announced on August 28, 2026 the addition of two evaluation sets to the Open ASR Leaderboard: Monsoon en-IN (Indian English) and Monsoon hi-IN (Hindi), according to a post on the Hugging Face blog. Hindi is described as the first Indic language on the leaderboard's multilingual tab, which the post says previously covered only European languages. Each language ships a public split available for self-scoring and a private split withheld to limit benchmark-specific optimisation. The four splits are speaker-disjoint and comprise 4,888 speakers, with 12 speaker attributes recorded per clip. According to the post, the sets were designed to vary along nine axes, including geography, age, gender, devices, and acoustic environments, in order to surface failure modes that an aggregate word error rate (WER) can hide. The blog cites prior research including 'Racial disparities in automated speech recognition,' which found commercial systems roughly twice as bad for Black speakers as for white speakers. The Indian English public set draws on 428 districts across 30 states and union territories and uses 315 to 582 distinct device models. Because Hindi has extensive spelling variation, the Hindi sets ship a lattice of accepted spellings per transcript span rather than a single reference string.
- Two new sets added: Monsoon en-IN and Monsoon hi-IN, published August 28, 2026
- 4,888 speakers across four speaker-disjoint splits, with 12 recorded speaker attributes
- Each language has a public split for self-scoring and a withheld private split
- Designed to vary along nine axes; Indian English set spans 428 districts and up to 582 device models
What it means for you
Speech-to-text tools are usually measured on English and European languages, which means they often work worse for accents and languages they were never tested on. This new benchmark tests Hindi and Indian English across thousands of real speakers, devices, and regions, giving a more honest picture of how well these tools actually work for those speakers. It's a step toward speech recognition that performs fairly for more of the world.
Try this
If your business handles Hindi or Indian-English voice input (call transcription, voice notes, dictation), check whether your speech-to-text vendor reports results on these Monsoon sets before assuming accuracy claims apply to your users.
Who should care
Teams building or buying voice transcription, call-center analytics, or voice assistants that serve Hindi or Indian-English speakers, and researchers working on multilingual ASR.
Skip this if
You don't work with speech recognition, or your users speak languages not covered here. This is a benchmark update, not a product you use directly.
Sources: Hugging Face — read the original