HomeDatasetshltcoe/megawika
M

hltcoe/megawika

Summarization · hltcoe· 14.6K
cc-by-sa-4.0 47 GB

MegaWika is a multi- and crosslingual text dataset containing 30 million Wikipedia passages with their scraped and cleaned web citations. The passages span 50 Wikipedias in 50 languages, and the articles in which the passages were originally embedded are included for convenience. Where a Wikipedia passage is in a non-English language, an automated English translation is provided. Furthermore, nearly 130 million English question/answer pairs were extracted from the passages, and FrameNet events occurring in the passages are detected using the LOME FrameNet parser.

Open in MLForge Sign up free Desktop app
# download instantly
mlforge datasets pull hltcoe/megawika

Dataset details

Task
Summarization
Language
af
License
cc-by-sa-4.0
Size
47 GB
Creator
hltcoe
Downloads
14.6K
Source
huggingface_datasets
Updated
2025-01-31

About hltcoe/megawika

- Homepage: HuggingFace - Repository: HuggingFace - Paper: [Coming soon] - Leaderboard: [Coming soon] - Point of Contact: Samuel Barham