Logo Lanfrica

yqdss13/XBMU-MC-A-Multilingual-Parallel-Corpus

Domaine:

natural language processing

Type de record:

dataset
Créateur:
yqd
Hôte:
In this study, a parallel corpus XBMU-MC (Northwest Minzu University-Multilingual Corpus) for multilingual machine translation and cross-language information retrieval tasks is constructed to address the relative scarcity of parallel corpora for low-resource languages (Chinese-Tibetan, Chinese-Western, Chinese-Mongolian). The original corpus consists of manually constructed specific bilingual texts and web-crawled publicly available multilingual data. The manually constructed corpus consists of Chinese texts written by team members in the fields of culture, science and technology, and society, and the target language translations are obtained through machine translation and manual proofreading; while the web-crawled corpus captures the original texts from mainstream Tibetan, Mongolian, and Viennese news websites to enlarge the scale of the corpus and the scope of the domains covered. After that, data enhancement techniques are used to expand the existing corpus, and the data are strictly screened to ensure the alignment consistency and translation accuracy between the source and target languages for each parallel corpus pair. Finally, 21,579 high-quality samples were screened and stored in JSON format, including three attributes: instruction, input and output.