HomeDatasetsmlfoundations/MINT-1T-HTML
M

mlfoundations/MINT-1T-HTML

Image To Text · mlfoundations· 147.8K
cc-by-4.0

🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-HTML.

Open in MLForge Sign up free Desktop app
# download instantly
mlforge datasets pull mlfoundations/MINT-1T-HTML

Dataset details

Task
Image To Text
Language
en
License
cc-by-4.0
Creator
mlfoundations
Downloads
147.8K
Source
huggingface_datasets
Updated
2024-09-21

About mlfoundations/MINT-1T-HTML

🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens