HomeDatasetsHuggingFaceFW/fineweb
F

HuggingFaceFW/fineweb

Text Generation · HuggingFaceFW· 421.7K
odc-by

🍷 FineWeb 15 trillion tokens of the finest data the 🌐 web has to offer What is it? The 🍷 FineWeb dataset consists of more than 18.5T tokens (originally 15T tokens) of cleaned and deduplicated english web data from CommonCrawl. The data processing pipeline is optimized for LLM performance and ran on the 🏭 datatrove library, our large scale data processing library. 🍷 FineWeb was originally meant to be a fully open replication of 🦅 RefinedWeb, with a… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb.

Open in MLForge Sign up free Desktop app
# download instantly
mlforge datasets pull HuggingFaceFW/fineweb

Dataset details

Task
Text Generation
Language
en
License
odc-by
Creator
HuggingFaceFW
Downloads
421.7K
Source
huggingface_datasets
Updated
2025-07-11

About HuggingFaceFW/fineweb

15 trillion tokens of the finest data the 🌐 web has to offer