HomeDatasetsLingoIITGN/PHINC
P

LingoIITGN/PHINC

Translation · LingoIITGN· 10.4K
cc-by-4.0 1.6 MB

Abstract Code-mixing is the phenomenon of using more than one language in a sentence. In the multilingual communities, it is a very frequently observed pattern of communication on social media platforms. Flexibility to use multiple languages in one text message might help to communicate efficiently with the target audience. But, the noisy user-generated code-mixed text adds to the challenge of pr

Open in MLForge Sign up free Desktop app
# download instantly
mlforge datasets pull LingoIITGN/PHINC

Dataset details

Task
Translation
Language
hi
License
cc-by-4.0
Size
1.6 MB
Rows / images
13.7K
Creator
LingoIITGN
Downloads
10.4K
Source
huggingface_datasets
Updated
2025-03-20

About LingoIITGN/PHINC

Abstract Code-mixing is the phenomenon of using more than one language in a sentence. In the multilingual communities, it is a very frequently observed pattern of communication on social media platforms. Flexibility to use multiple languages in one text message might help to communicate efficiently with the target audience. But, the noisy user-generated code-mixed text adds to the challenge of processing and understanding natural language to a much larger extent. Machine translation from monolingual source to the target language is a well-studied research problem. Here, we demonstrate that widely popular and sophisticated translation systems such as Google Translate fail at times to translate code-mixed text effectively. To address this challenge, we present a parallel corpus of the 13,738 code-mixed Hindi-English sentences and their corresponding human translation in English. In addition, we also propose a translation pipeline build on top of Google Translate. The evaluation of the proposed pipeline on PHINC demonstrates an increase in the performance of the underlying system. With minimal effort, we can extend the dataset and the proposed approach to other code-mixing language pairs.