HomeDatasetsNCSpeech/YO-CPT-ru
Y

NCSpeech/YO-CPT-ru

Text To Speech · NCSpeech· 6.9K
other 940 GB

YouTube-Oriented dataset for Continual Pre-Training (Russian). A large, heavily quality-filtered corpus of Russian speech mined from YouTube (via YODAS2) and processed into clean, single-speaker, TTS-grade utterances. Every utterance ships with an ensemble-verified transcription, a punctuated/denormalized and stress-marked text variant, word-level forced alignment, within- and cross-video speaker

Open in MLForge Sign up free Desktop app
# download instantly
mlforge datasets pull NCSpeech/YO-CPT-ru

Dataset details

Task
Text To Speech
Language
ru
License
other
Size
940 GB
Rows / images
1.6M
Creator
NCSpeech
Downloads
6.9K
Source
huggingface_datasets
Updated
2026-08-04

About NCSpeech/YO-CPT-ru

YouTube-Oriented dataset for Continual Pre-Training (Russian). A large, heavily quality-filtered corpus of Russian speech mined from YouTube (via YODAS2) and processed into clean, single-speaker, TTS-grade utterances. Every utterance ships with an ensemble-verified transcription, a punctuated/denormalized and stress-marked text variant, word-level forced alignment, within- and cross-video speaker identities, an audio-quality (MOS) score, and a speaker persona built from the voice and, where available, the speaker's on-screen face.