HomeDatasetsnvidia/OCR-Synthetic-Multilingual-v1
O

nvidia/OCR-Synthetic-Multilingual-v1

Object Detection · nvidia· 8.7K
cc-by-4.0 47 GB

OCR-Synthetic-Multilingual-v1 Dataset Description Large-scale synthetically generated OCR training dataset for multilingual text detection and recognition. The data was produced using a heavily modified and extended version of SynthDoG (Synthetic Document Generator), originally introduced in the Donut project by Kim et al. This dataset was used to train Nemotron OCR v2, a state-of-the-art multilingual OCR model that is part of the NVIDIA NeMo Retriever collection.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/OCR-Synthetic-Multilingual-v1.

Open in MLForge Sign up free Desktop app
# download instantly
mlforge datasets pull nvidia/OCR-Synthetic-Multilingual-v1

Dataset details

Task
Object Detection
Language
en
License
cc-by-4.0
Size
47 GB
Creator
nvidia
Downloads
8.7K
Source
huggingface_datasets
Updated
2026-04-20

About nvidia/OCR-Synthetic-Multilingual-v1

Large-scale synthetically generated OCR training dataset for multilingual text detection and recognition. The data was produced using a heavily modified and extended version of SynthDoG (Synthetic Document Generator), originally introduced in the Donut project by Kim et al.