HomeDatasetsfacebook/voxpopuli
V

facebook/voxpopuli

Automatic Speech Recognition · facebook· 31.7K
["cc0-1.0","other"]

Dataset Card for Voxpopuli Dataset Summary VoxPopuli is a large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation. The raw data is collected from 2009-2020 European Parliament event recordings. We acknowledge the European Parliament for creating and sharing these materials. This implementation contains transcribed speech data for 18 languages. It also contains 29 hours of transcribed speech data of non-native… See the full description on the dataset page: https://huggingface.co/datasets/facebook/voxpopuli.

Open in MLForge Sign up free Desktop app
# download instantly
mlforge datasets pull facebook/voxpopuli

Dataset details

Task
Automatic Speech Recognition
Language
en
License
["cc0-1.0","other"]
Creator
facebook
Downloads
31.7K
Source
huggingface_datasets
Updated
2026-01-30

About facebook/voxpopuli

Table of Contents - Table of Contents - Dataset Description - Dataset Summary - Supported Tasks and Leaderboards - Languages - Dataset Structure - Data Instances - Data Fields - Data Splits - Dataset Creation - Curation Rationale - Source Data - Annotations - Personal and Sensitive Information - Considerations for Using the Data - Social Impact of Dataset - Discussion of Biases - Other Known Limitations - Additional Information - Dataset Curators - Licensing Information - Citation Information - Contributions