HomeDatasetsCohereLabs/aya_dataset
A

CohereLabs/aya_dataset

Other · CohereLabs· 17.4K
apache-2.0 477 MB

Dataset Summary The Aya Dataset is a multilingual instruction fine-tuning dataset curated by an open-science community via Aya Annotation Platform from Cohere Labs. The dataset contains a total of 204k human-annotated prompt-completion pairs along with the demographics data of the annotators. This dataset can be used to train, finetune, and evaluate multilingual LLMs. Curated by: Contributors of Aya Open Science Intiative. Language(s): 65 languages (71 including dialects &… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/aya_dataset.

Open in MLForge Sign up free Desktop app
# download instantly
mlforge datasets pull CohereLabs/aya_dataset

Dataset details

Task
Other
Language
amh
License
apache-2.0
Size
477 MB
Creator
CohereLabs
Downloads
17.4K
Source
huggingface_datasets
Updated
2025-04-15

About CohereLabs/aya_dataset

Dataset Summary The Aya Dataset is a multilingual instruction fine-tuning dataset curated by an open-science community via Aya Annotation Platform from Cohere Labs. The dataset contains a total of 204k human-annotated prompt-completion pairs along with the demographics data of the annotators. This dataset can be used to train, finetune, and evaluate multilingual LLMs.