Datasets

Dataset

indicvoices_r

by ai4bharat

IndicVoices-R: Multilingual, Multi-Speaker Speech Corpus for Indian TTS Dataset Summary IndicVoices-R (IV-R) is the largest multilingual Indian text-to-speech (TTS) dataset derived from an automatic speech recognition (ASR) dataset. It contains 1,704 hours …

CC-BY-4.0text-to-speech 7666 35100K < n < 1M rows

published 4 Mar 2025

Preview

The dataset does not exist, or is not accessible without authentication (private or gated). Please check the spelling of the dataset name or retry with authentication.

Cite this

CC-BY-4.0audiotextCite

Tags

task_categories:text-to-speechlanguage:aslanguage:bnlanguage:gulanguage:hilanguage:knlanguage:kslanguage:mllanguage:mrlanguage:nelanguage:orlanguage:palanguage:salanguage:talanguage:telanguage:urlicense:cc-by-4.0size_categories:100K<n<1Mformat:parquetmodality:audiomodality:textlibrary:datasetslibrary:dasklibrary:polarslibrary:mlcroissantregion:us