HomeDatasetsOpenGVLab/OmniCorpus-CC-210M
O

OpenGVLab/OmniCorpus-CC-210M

Image To Text · OpenGVLab· 21.5K
cc-by-4.0 450 GB

🐳 OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text This repository contains 210 million image-text interleaved documents filtered from the OmniCorpus-CC dataset, which was sourced from Common Crawl. Repository: https://github.com/OpenGVLab/OmniCorpus Paper (ICLR 2025 Spotlight): https://arxiv.org/abs/2406.08418 OmniCorpus dataset is a large-scale image-text interleaved dataset, which pushes the boundaries of scale and diversity by encompassing… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/OmniCorpus-CC-210M.

Open in MLForge Sign up free Desktop app
# download instantly
mlforge datasets pull OpenGVLab/OmniCorpus-CC-210M

Dataset details

Task
Image To Text
Language
en
License
cc-by-4.0
Size
450 GB
Creator
OpenGVLab
Downloads
21.5K
Source
huggingface_datasets
Updated
2025-03-20