HomeDatasetsSanghyang00/omniasr-molge
O

Sanghyang00/omniasr-molge

Automatic Speech Recognition · Sanghyang00· 30.3K
other

Training-friendly re-segmentation of Meta’s facebook/omnilingual-asr-corpus: long utterances are segmented / aligned into ≤30s clips with transcripts, then packed as Parquet shards with embedded FLAC.

Open in MLForge Sign up free Desktop app
# download instantly
mlforge datasets pull Sanghyang00/omniasr-molge

Dataset details

Task
Automatic Speech Recognition
License
other
Rows / images
2.6M
Creator
Sanghyang00
Downloads
30.3K
Source
huggingface_datasets
Updated
2026-08-04

About Sanghyang00/omniasr-molge

Training-friendly re-segmentation of Meta’s facebook/omnilingual-asr-corpus: long utterances are segmented / aligned into ≤30s clips with transcripts, then packed as Parquet shards with embedded FLAC.