HomeDatasetsgoogle/WaxalNLP
W

google/WaxalNLP

Automatic Speech Recognition · google· 62.6K
["cc-by-sa-4.0","cc-by-4.0"]

Waxal Datasets The WAXAL dataset is a large-scale multilingual speech corpus for African languages, introduced in the paper WAXAL: A Large-Scale Multilingual African Language Speech Corpus. Dataset Description The Waxal project provides datasets for both Automated Speech Recognition (ASR) and Text-to-Speech (TTS) for African languages. The goal of this dataset's creation and release is to facilitate research that improves the accuracy and fluency of speech and… See the full description on the dataset page: https://huggingface.co/datasets/google/WaxalNLP.

Open in MLForge Sign up free Desktop app
# download instantly
mlforge datasets pull google/WaxalNLP

Dataset details

Task
Automatic Speech Recognition
Language
ach
License
["cc-by-sa-4.0","cc-by-4.0"]
Creator
google
Downloads
62.6K
Source
huggingface_datasets
Updated
2026-08-02

About google/WaxalNLP

The WAXAL dataset is a large-scale multilingual speech corpus for African languages, introduced in the paper WAXAL: A Large-Scale Multilingual African Language Speech Corpus.