HomeDatasetsMutonix/Vript
V

Mutonix/Vript

Video Classification · Mutonix· 15.7K
Unknown 477 MB

🎬 Vript: Refine Video Captioning into Video Scripting [Github Repo] We construct a fine-grained video-text dataset with 12K annotated high-resolution videos (~400k clips). The annotation of this dataset is inspired by the video script. If we want to make a video, we have to first write a script to organize how to shoot the scenes in the videos. To shoot a scene, we need to decide the content, shot type (medium shot, close-up, etc), and how the camera moves (panning, tilting, etc).… See the full description on the dataset page: https://huggingface.co/datasets/Mutonix/Vript.

Open in MLForge Sign up free Desktop app
# download instantly
mlforge datasets pull Mutonix/Vript

Dataset details

Task
Video Classification
Language
en
License
Unknown
Size
477 MB
Creator
Mutonix
Downloads
15.7K
Source
huggingface_datasets
Updated
2024-06-11

About Mutonix/Vript

🎬 Vript: Refine Video Captioning into Video Scripting [Github Repo] --- We construct a fine-grained video-text dataset with 12K annotated high-resolution videos (~400k clips). The annotation of this dataset is inspired by the video script. If we want to make a video, we have to first write a script to organize how to shoot the scenes in the videos. To shoot a scene, we need to decide the content, shot type (medium shot, close-up, etc), and how the camera moves (panning, tilting, etc). Therefore, we extend video captioning to video scripting by annotating the videos in the format of video scripts. Different from the previous video-text datasets, we densely annotate the entire videos without discarding any scenes and each scene has a caption with ~145 words. Besides the vision modality, we transcribe the voice-over into text and put it along with the video title to give more background information for annotating the videos.