Datasets

Dataset

Vaani

by ARTPARK-IISc

VAANI is an India-representative multi-modal multi-lingual dataset. The current version (phase 1- 80 districts, phase 2- 85 districts) contains ~31,255 hours of spontaenous,image-prompted speech by 156K speakers across 165 districts, talking about 288K imag…

CC-BY-4.0automatic-speech-recognition 19212 14410M < n < 100M rows

published 30 Sept 2024

Preview

The dataset does not exist, or is not accessible without authentication (private or gated). Please check the spelling of the dataset name or retry with authentication.

Cite this

CC-BY-4.0Cite

Tags

task_categories:automatic-speech-recognitiontask_categories:text-to-speechtask_categories:image-to-texttask_categories:text-to-imagelanguage:nelanguage:aslanguage:mllanguage:gulanguage:orlanguage:enlanguage:talanguage:urlanguage:telanguage:knlanguage:bnlanguage:hilicense:cc-by-4.0size_categories:10M<n<100Marxiv:2603.28714region:us