Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Challenge Track: Divide and Translate: Parameter Isolation with Encoder Freezing for Low-Resource Indic NM

Domain:

natural language processing

Record type:

paper
Creator:
AssKan
Publisher:
Und
Host:avatar
Low-resource tribal languages remain severely underrepresented in modern machine translation, especially when training data is small, noisy, and linguistically diverse. We propose Divide and Translate, a modular framework that replaces unified multilingual fine-tuning with direction-specific expert models. A single frozen NLLB encoder provides a stable multilingual representation space, while independent decoders specialize in individual translation directions. To improve robustness, we apply bitext reversal augmentation, doubling supervision and enabling balanced bi-directional learning without generating synthetic data. This design reduces gradient conflict, mitigates hallucination, and improves generalization under domain shift. Despite minimal compute and limited data, our system achieves strong leaderboard performance with low validation-test divergence, showing that specialization and parameter isolation outperform scale for under-resourced languages.

Visit

doi.org

Tasks

machine translation

Tags

Computational LinguisticsArtificial Intelligence

Similar

IndiAnn: An Annotation Platform for Low-Resource Indic LanguagesParameter Scaling in Encoder-Only Models for Zero-Shot Cross-Lingual Accuracy on XTREME-R Low-Resource LanguagesImproving Multilingual Semantic Textual Similarity with Shared Sentence Encoder for Low-resource LanguagesRobust Multilingual Encoder Training for Low-Resource Language AlignmentMOHAMEDFAIZN/digital-divide-low-resource-aiExploring Graph-based Transformer Encoder for Low-Resource Neural Machine Translation

IndiAnn: An Annotation Platform for Low-Resource Indic Languages

Linguistic annotation tools that work well for non-Indic languages (e.g. English, German, Spanish, e

Parameter Scaling in Encoder-Only Models for Zero-Shot Cross-Lingual Accuracy on XTREME-R Low-Resource Languages

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Improving Multilingual Semantic Textual Similarity with Shared Sentence Encoder for Low-resource Languages

Measuring the semantic similarity between two sentences (or Semantic Textual Similarity - STS) is fu

Robust Multilingual Encoder Training for Low-Resource Language Alignment

Pre-trained multilingual language encoders, such as multilingual BERT and XLM-R, show great potentia

MOHAMEDFAIZN/digital-divide-low-resource-ai

Offline AI education platform for low-resource learners using Ollama, WebLLM, React & Edge AI. # 🌐

Exploring Graph-based Transformer Encoder for Low-Resource Neural Machine Translation

The Transformer is commonly used in Neural Machine Translation (NMT), but it faces issues with over-