RESPIN-S1.0: A read speech corpus of 10000+ hours in dialects of nine Indian Languages

Saurabh Kumar (Indian Institute of Science, Bangalore) · Abhayjeet Singh (Indian Institute of Science, Indian institute of science, Bangalore) · DEEKSHITHA G (Indian Institute of Science, Indian institute of science, Bangalore) · Amartya veer (Indian Institute of Science) · Jesuraj Bandekar (Indian Institute of Science) · Savitha Murthy (Indian Institute of Science, Indian institute of science, Bangalore) · Sumit Sharma (Indian Institute of Science, Indian institute of science, Bangalore) · Sandhya Badiger (German Research Center for AI) · Sathvik Udupa (Brno University of Technology) · Amala Nagireddi (Indian Institute of Science, Indian institute of science, Bangalore) · Srinivasa Raghavan K M (Navana Tech India Private Limited) · Rohan Saxena (Navanatech) · Jai Nanavati (Cornell University) · Raoul Nanavati (Navana Tech India Private LImited) · Janani Sridharan (Navana.ai) · Arjun Mehta (HireGit) · Ashish S (HiFX IT & Media Services) · Sai Mora (S2T.AI) · Prashanthi Venkataramakrishnan (Turing) · Gauri Date (TransPerfect) · Karthika P (Navana Tech India Private Limited) · Prasanta Ghosh (Indian Institute of Science, Bangalore)
asr modelsbenchmark performancecrowdsourced platformdialect-rich corpusdialectal variatione-branchformerindian languageslanguage identificationlow-resource settingsmultilingual asrphonetic lexiconsrespin-s1.0self-supervised modelsspeech recognitiontdnn-hmmtranscription quality

We introduce **RESPIN-S1.0**, the largest publicly available dialect-rich read-speech corpus for Indian languages, comprising more than 10,000 hours of validated audio across nine major languages: Bengali, Bhojpuri, Chhattisgarhi, Hindi, Kannada, Magahi, Maithili, Marathi, and Telugu. Indian languages exhibit high dialectal variation and are spoken by populations that remain digitally underserved. Existing speech corpora typically represent only standard dialects and lack domain and linguistic diversity. RESPIN-S1.0 addresses this limitation by collecting speech across more than 38 dialects and two high-impact domains: agriculture and finance. Text data were composed by native dialect speakers and validated through a pipeline combining automated and manual checks. Over 200,000 unique sentences were recorded through a crowdsourced mobile platform and categorised into clean, semi-noisy, and noisy subsets based on transcription quality, with the clean portion alone exceeding 10,000 hours. Along with audio and transcriptions, RESPIN provides dialect-aware phonetic lexicons, speaker metadata, and reproducible train, development, and test splits. To benchmark performance, we evaluate multiple ASR models, including TDNN-HMM, E-Branchformer, Whisper, and wav2vec2-based self-supervised models, and find that fine-tuning on RESPIN significantly improves recognition accuracy over pretrained baselines. A subset of RESPIN-S1.0 has already supported community challenges such as the SLT Code Hackathon 2022 and MADASR@ASRU 2023 and 2025, releasing more than 1,200 hours publicly. This resource supports research in dialectal ASR, language identification, and related speech technologies, establishing a comprehensive benchmark for inclusive, dialect-rich ASR in multilingual low-resource settings. **Dataset:** https://spiredatasets.ee.iisc.ac.in/respincorpus **Code:** https://github.com/labspire/respin_baselines.git