HomeDatasetsMLCommons/peoples_speech
P

MLCommons/peoples_speech

Automatic Speech Recognition · MLCommons· 66.2K
["cc-by-2.0","cc-by-2.5","cc-by-3.0","cc-by-4.0","cc-by-sa-3.0","cc-by-sa-4.0"]

Dataset Card for People's Speech Dataset Summary The People's Speech Dataset is among the world's largest English speech recognition corpus today that is licensed for academic and commercial usage under CC-BY-SA and CC-BY 4.0. It includes 30,000+ hours of transcribed speech in English languages with a diverse set of speakers. This open dataset is large enough to train speech-to-text systems and crucially is available with a permissive license. Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/peoples_speech.

Open in MLForge Sign up free Desktop app
# download instantly
mlforge datasets pull MLCommons/peoples_speech

Dataset details

Task
Automatic Speech Recognition
Language
en
License
["cc-by-2.0","cc-by-2.5","cc-by-3.0","cc-by-4.0","cc-by-sa-3.0","cc-by-sa-4.0"]
Creator
MLCommons
Downloads
66.2K
Source
huggingface_datasets
Updated
2024-11-20

About MLCommons/peoples_speech

Table of Contents - Dataset Description - Dataset Summary - Supported Tasks - Languages - Dataset Structure - Data Instances - Data Fields - Data Splits - Dataset Creation - Curation Rationale - Source Data - Annotations - Personal and Sensitive Information - Considerations for Using the Data - Social Impact of Dataset - Discussion of Biases - Other Known Limitations - Additional Information - Dataset Curators - Licensing Information - Citation Information