HomeDatasetsHuggingFaceM4/FineVisionMax
F

HuggingFaceM4/FineVisionMax

Image Text To Text · HuggingFaceM4· 91.3K
Unknown 47 GB

Fine Vision FineVision is a massive collection of datasets with 17.3M images, 24.3M samples, 88.9M turns, and 9.5B answer tokens, designed for training state-of-the-art open Vision-Language-Models. More detail can be found in the blog post: https://huggingface.co/spaces/HuggingFaceM4/FineVision The version in this repository concatenated all the configs in the original dataset and then shuffled them. This is done to facilitate streaming the data directly from the hub! Load… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceM4/FineVisionMax.

Open in MLForge Sign up free Desktop app
# download instantly
mlforge datasets pull HuggingFaceM4/FineVisionMax

Dataset details

Task
Image Text To Text
Language
en
License
Unknown
Size
47 GB
Creator
HuggingFaceM4
Downloads
91.3K
Source
huggingface_datasets
Updated
2025-10-21

About HuggingFaceM4/FineVisionMax

FineVision is a massive collection of datasets with 17.3M images, 24.3M samples, 88.9M turns, and 9.5B answer tokens, designed for training state-of-the-art open Vision-Language-Models.