Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Low-Resource English-Tigrinya MT: Leveraging Multilingual Models, Custom Tokenizers, and Clean Evaluation Benchmarks

Domain:

natural language processing

Record type:

paperdatasetmodel
Creator:
TekGidNej
Host:avatar
Despite advances in Neural Machine Translation (NMT), low-resource languages like Tigrinya remain underserved due to persistent challenges, including limited corpora, inadequate tokenization strategies, and the lack of standardized evaluation benchmarks. This paper investigates transfer learning techniques using multilingual pretrained models to enhance translation quality for morphologically rich, low-resource languages. We propose a refined approach that integrates language-specific tokenization, informed embedding initialization, and domain-adaptive fine-tuning. To enable rigorous assessment, we construct a high-quality, human-aligned English-Tigrinya evaluation dataset covering diverse domains. Experimental results demonstrate that transfer learning with a custom tokenizer substantially outperforms zero-shot baselines, with gains validated by BLEU, chrF, and qualitative human evaluation. Bonferroni correction is applied to ensure statistical significance across configurations. Error analysis reveals key limitations and informs targeted refinements. This study underscores the importance of linguistically aware modeling and reproducible benchmarks in bridging the performance gap for underrepresented languages. Resources are available at github.com and huggingface.co This submission is 8 pages long, includes 4 tables, and contains all required conference details

Visit

arxiv.org

Tasks

machine translation

Languages

Tigrigna

Tags

Computation and LanguageArtificial Intelligence68T50, 68T35I.2.7; H.3.1; I.2.6

Similar

nuredinali/Tigrinya-English-MT-Evaluation-DatasetsAdvancing Fact-Checking in Low-Resource and Multilingual Contexts: Benchmarks, Frameworks, and EvaluationsScaling Pretrained Models and Intermediate-Task Training in Low-Resource Cross-Lingual BenchmarksLeveraging LLMs for MT in Crisis Scenarios: a blueprint for low-resource languagestyracs18/low-resource-mtA Low-Resource English–Hassaniya Parallel Corpus with Neural Machine Translation Benchmarks

nuredinali/Tigrinya-English-MT-Evaluation-Datasets

**About**: This is the dataset for the project “Error Analysis of Tigrinya-English Machine Translat

Advancing Fact-Checking in Low-Resource and Multilingual Contexts: Benchmarks, Frameworks, and Evaluations

The growing reliance on Large Language Models (LLMs) and social media platforms for information cons

Scaling Pretrained Models and Intermediate-Task Training in Low-Resource Cross-Lingual Benchmarks

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Leveraging LLMs for MT in Crisis Scenarios: a blueprint for low-resource languages

In an evolving landscape of crisis communication, the need for robust and adaptable Machine Translat

tyracs18/low-resource-mt

# In-Context Learning vs. Pivot-Based Translation for Low-Resource MT #### Overview This project co

A Low-Resource English–Hassaniya Parallel Corpus with Neural Machine Translation Benchmarks