Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

A Case Against Implicit Standards: Homophone Normalization in Machine Translation for Languages that use the Ge'ez Script

Domain:

natural language processing

Record type:

paper
Creator:
NigTonja, Atnafu LambeboAdeAle
Host:avatar
Homophone normalization, where characters that have the same sound in a writing script are mapped to one character, is a pre-processing step applied in Amharic Natural Language Processing (NLP) literature. While this may improve performance reported by automatic metrics, it also results in models that are not able to understand different forms of writing in a single language. Further, there might be impacts in transfer learning, where models trained on normalized data do not generalize well to other languages. In this paper, we experiment with monolingual training and cross-lingual transfer to understand the impacts of normalization on languages that use the Ge'ez script. We then propose a post-inference intervention in which normalization is applied to model predictions instead of training data. With our simple scheme of post-inference normalization, we show that we can achieve an increase in BLEU score of up to 1.03 while preserving language features in training. Our work contributes to the broader discussion on technology-facilitated language change and calls for more language-aware interventions. Paper under review

Visit

arxiv.org

Tasks

machine translationtext normalizationtransfer learning

Languages

Amharic

Tags

Computation and LanguageArtificial Intelligence

Similar

Machine Translation for Ge'ez LanguageImpacts of Homophone Normalization on Semantic Models for AmhariImpacts of Homophone Normalization on Semantic Models for AmharicScript Normalization for Unconventional Writing of Under-Resourced Languages in Bilingual CommunitiesRecovering Whisper's Wasted Output Capacity for Ge'ez-Script Languages: A Controlled Tokenizer-Extension Study on AmharicThe Effect of Normalization for Bi-directional Amharic-English Neural Machine Translation

Machine Translation for Ge'ez Language

Machine translation (MT) for low-resource languages such as Ge'ez, an ancient language that is no lo

Impacts of Homophone Normalization on Semantic Models for Amhari

Amharic is the second-most spoken Semitic language after Arabic and serves as the official working language of the government of Ethiopia. In Amharic writing, there are different characters with the same sound, which are called homophones. The current trend in Amha

Impacts of Homophone Normalization on Semantic Models for Amharic

Script Normalization for Unconventional Writing of Under-Resourced Languages in Bilingual Communities

The wide accessibility of social media has provided linguistically under-represented communities wit

Recovering Whisper's Wasted Output Capacity for Ge'ez-Script Languages: A Controlled Tokenizer-Extension Study on Amharic

The Effect of Normalization for Bi-directional Amharic-English Neural Machine Translation

Machine translation (MT) is one of the main tasks in natural language processing whose objective is