Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

A PARALLEL CORPUS FOR ADVANCING ENGLISH–SANTALI NEURAL MACHINE TRANSLATION

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Sun
Éditeur:
Aca
Hôte:
Machine Translation (MT) poses a significant challenge in developing language corpora for low-resource languages due to their minimal digital availability. Building such corpora is essential for preserving and promoting these languages. Santali, for instance, has very limited representation across online resources, and no proper translation tools including Google Translate exist for it. Developing a translation framework under such constraints is particularly difficult, as issues like low translation accuracy and heavy computational requirements arise. To overcome these limitations, the proposed MT system employs EnSanCorp, an English-Santali parallel corpus designed to facilitate Neural Machine Translation (NMT). EnSanCorp is created using multiple approaches, such as web-based parallel data extraction and optical character recognition (OCR) applied to scanned documents. The OCR-based method also demonstrates its usefulness for building corpora of other low-resource languages lacking online data. EnSanCorp currently contains 5,930 aligned sentences, 39,646 English tokens, and 39,936 Santali tokens, making it the most extensive English-Santali corpus available for research and non-commercial purposes. Evaluation results show that the Bilingual Evaluation Understudy (BLEU) scores for Statistical Machine Translation (SMT) and NMT vary across word and sentence levels: for word pairs, the scores are 0.04 (SMT) and 1.10 (NMT); for sentence pairs, 1.15 (SMT) and 7.20 (NMT). The overall BLEU scores achieved are 0.05 for SMT and 3.10 for NMT.

Visit

doi.org

Tasks

machine translation