HomeDatasetsneh7777/Pretraining-V1
P

neh7777/Pretraining-V1

Text To Speech · neh7777· 5.1K
cc-by-4.0 7.2 TB

A large-scale, unified collection of speech data for text-to-speech (TTS) and speech research. This dataset consolidates 17 distinct source datasets into a single, schema-normalized resource covering Indian / South Asian languages, plus major European, African, MENA, and Central Asian languages, with over 13.7 million utterances and 26,000+ hours of audio.

Open in MLForge Sign up free Desktop app
# download instantly
mlforge datasets pull neh7777/Pretraining-V1

Dataset details

Task
Text To Speech
Language
hi
License
cc-by-4.0
Size
7.2 TB
Rows / images
16.2M
Creator
neh7777
Downloads
5.1K
Source
huggingface_datasets
Updated
2026-05-15

About neh7777/Pretraining-V1

A large-scale, unified collection of speech data for text-to-speech (TTS) and speech research. This dataset consolidates 17 distinct source datasets into a single, schema-normalized resource covering Indian / South Asian languages, plus major European, African, MENA, and Central Asian languages, with over 13.7 million utterances and 26,000+ hours of audio.