HomeDatasetsVisGym/visgym_data
V

VisGym/visgym_data

Image Text To Text · VisGym· 12.5K
apache-2.0 4.7 GB

VisGym Dataset Project Page | Paper | GitHub VisGym consists of 17 diverse, long-horizon environments designed to systematically evaluate, diagnose, and train Vision-Language Models (VLMs) on visually interactive tasks. In these environments, agents must select actions conditioned on both their past actions and observation history, challenging their ability to handle complex, multimodal sequences. Dataset Summary This dataset contains trajectories and interaction data… See the full description on the dataset page: https://huggingface.co/datasets/VisGym/visgym_data.

Open in MLForge Sign up free Desktop app
# download instantly
mlforge datasets pull VisGym/visgym_data

Dataset details

Task
Image Text To Text
Language
en
License
apache-2.0
Size
4.7 GB
Creator
VisGym
Downloads
12.5K
Source
huggingface_datasets
Updated
2026-02-05

About VisGym/visgym_data

VisGym consists of 17 diverse, long-horizon environments designed to systematically evaluate, diagnose, and train Vision-Language Models (VLMs) on visually interactive tasks. In these environments, agents must select actions conditioned on both their past actions and observation history, challenging their ability to handle complex, multimodal sequences.