HomeDatasetslegacy-datasets/wikipedia
W

legacy-datasets/wikipedia

Text Generation · legacy-datasets· 110.7K
["cc-by-sa-3.0","gfdl"] 488 KB

Wikipedia dataset containing cleaned articles of all languages. The datasets are built from the Wikipedia dump (https://dumps.wikimedia.org/) with one split per language. Each example contains the content of one full Wikipedia article with cleaning to strip markdown and unwanted sections (references, etc.).

Open in MLForge Sign up free Desktop app
# download instantly
mlforge datasets pull legacy-datasets/wikipedia

Dataset details

Task
Text Generation
Language
aa
License
["cc-by-sa-3.0","gfdl"]
Size
488 KB
Creator
legacy-datasets
Downloads
110.7K
Source
huggingface_datasets
Updated
2024-03-11

About legacy-datasets/wikipedia

Table of Contents - Dataset Description - Dataset Summary - Supported Tasks and Leaderboards - Languages - Dataset Structure - Data Instances - Data Fields - Data Splits - Dataset Creation - Curation Rationale - Source Data - Annotations - Personal and Sensitive Information - Considerations for Using the Data - Social Impact of Dataset - Discussion of Biases - Other Known Limitations - Additional Information - Dataset Curators - Licensing Information - Citation Information - Contributions