HomeDatasetsopenslr/librispeech_asr
L

openslr/librispeech_asr

Automatic Speech Recognition · openslr· 58.4K
["cc-by-4.0"] 477 MB

Dataset Card for librispeech_asr Dataset Summary LibriSpeech is a corpus of approximately 1000 hours of 16kHz read English speech, prepared by Vassil Panayotov with the assistance of Daniel Povey. The data is derived from read audiobooks from the LibriVox project, and has been carefully segmented and aligned. Supported Tasks and Leaderboards automatic-speech-recognition, audio-speaker-identification: The dataset can be used to train a model for Automatic… See the full description on the dataset page: https://huggingface.co/datasets/openslr/librispeech_asr.

Open in MLForge Sign up free Desktop app
# download instantly
mlforge datasets pull openslr/librispeech_asr

Dataset details

Task
Automatic Speech Recognition
Language
en
License
["cc-by-4.0"]
Size
477 MB
Creator
openslr
Downloads
58.4K
Source
huggingface_datasets
Updated
2025-07-25

About openslr/librispeech_asr

Table of Contents - Dataset Description - Dataset Summary - Supported Tasks and Leaderboards - Languages - Dataset Structure - Data Instances - Data Fields - Data Splits - Dataset Creation - Curation Rationale - Source Data - Annotations - Personal and Sensitive Information - Considerations for Using the Data - Social Impact of Dataset - Discussion of Biases - Other Known Limitations - Additional Information - Dataset Curators - Licensing Information - Citation Information - Contributions