Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

A PARALLEL CORPUS FOR ADVANCING ENGLISH–SANTALI NEURAL MACHINE TRANSLATION

Domain:

natural language processing

Record type:

dataset
Creator:
Sun
Publisher:
Aca
Host:
Machine Translation (MT) poses a significant challenge in developing language corpora for low-resource languages due to their minimal digital availability. Building such corpora is essential for preserving and promoting these languages. Santali, for instance, has very limited representation across online resources, and no proper translation tools including Google Translate exist for it. Developing a translation framework under such constraints is particularly difficult, as issues like low translation accuracy and heavy computational requirements arise. To overcome these limitations, the proposed MT system employs EnSanCorp, an English-Santali parallel corpus designed to facilitate Neural Machine Translation (NMT). EnSanCorp is created using multiple approaches, such as web-based parallel data extraction and optical character recognition (OCR) applied to scanned documents. The OCR-based method also demonstrates its usefulness for building corpora of other low-resource languages lacking online data. EnSanCorp currently contains 5,930 aligned sentences, 39,646 English tokens, and 39,936 Santali tokens, making it the most extensive English-Santali corpus available for research and non-commercial purposes. Evaluation results show that the Bilingual Evaluation Understudy (BLEU) scores for Statistical Machine Translation (SMT) and NMT vary across word and sentence levels: for word pairs, the scores are 0.04 (SMT) and 1.10 (NMT); for sentence pairs, 1.15 (SMT) and 7.20 (NMT). The overall BLEU scores achieved are 0.05 for SMT and 3.10 for NMT.

Visit

doi.org

Tasks

machine translation

Similar

Amharic-English Parallel Corpus for Neural Machine TranslationEnglish-Twi Parallel Corpus for Machine TranslationA Low-Resource English–Hassaniya Parallel Corpus with Neural Machine Translation BenchmarksCrowdsourcing Parallel Corpus for English-Oromo Neural Machine Translation using Community Engagement PlatformExtended Parallel Corpus for Amharic-English Machine TranslationAmhEn: Amharic-English Large Parallel Corpus for Machine Translation

Amharic-English Parallel Corpus for Neural Machine Translation

Amharic is the working language of Ethiopia and, owing to its Semitic characteristics, the language

English-Twi Parallel Corpus for Machine Translation

We present a parallel machine translation training corpus for English and Akuapem Twi of 25,421 sent

A Low-Resource English–Hassaniya Parallel Corpus with Neural Machine Translation Benchmarks

Crowdsourcing Parallel Corpus for English-Oromo Neural Machine Translation using Community Engagement Platform

Even though Afaan Oromo is the most widely spoken language in the Cushitic family by more than fifty million people in the Horn and East Africa, it is surprisingly resource-scarce from a technological point of view. The increasing amount of various useful documents

Extended Parallel Corpus for Amharic-English Machine Translation

This paper describes the acquisition, preprocessing, segmentation, and alignment of an Amharic-English parallel corpus. It will be helpful for machine translation of a low-resource language, Amharic. We freely released the corpus for research purposes. Furthermore,

AmhEn: Amharic-English Large Parallel Corpus for Machine Translation

Recently, using deep neural networks for machine translation (MT) tasks has received great attention. In order for these networks to learn abstract representations of the input and store them as continuous vectors, they need a lot of data. However, very few researc