Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Script Normalization for Unconventional Writing of Under-Resourced Languages in Bilingual Communities

Domaine:

natural language processing

Type de record:

paper
Créateur:
AhmAnastasopoulos, Antonios
Hôte:avatar
The wide accessibility of social media has provided linguistically under-represented communities with an extraordinary opportunity to create content in their native languages. This, however, comes with certain challenges in script normalization, particularly where the speakers of a language in a bilingual community rely on another script or orthography to write their native language. This paper addresses the problem of script normalization for several such languages that are mainly written in a Perso-Arabic script. Using synthetic data with various levels of noise and a transformer-based model, we demonstrate that the problem can be effectively remediated. We conduct a small-scale evaluation of real data as well. Our experiments indicate that script normalization is also beneficial to improve the performance of downstream tasks such as machine translation and language identification. To appear in the proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL)

Visit

arxiv.org

Tasks

text normalization

Tags

Computation and Language

Similaires

Multilingualism in Under-resourced Languages for Sustainable Development in Rural CommunitiesDatasheets for Under-resourced Languages: An ExampleShort Text Language Identification for Under Resourced LanguagesStrategies for building wordnets for under-resourced languages: The case of African languagesA Case Against Implicit Standards: Homophone Normalization in Machine Translation for Languages that use the Ge'ez ScriptTranslation-Based Dictionary Alignment for Under-Resourced Bantu Languages

Multilingualism in Under-resourced Languages for Sustainable Development in Rural Communities

Cameroon, a central African country, is one of the most linguistically diverse countries in Africa w

Datasheets for Under-resourced Languages: An Example

The datasheet provides an example of how to use the Datasheet standard for describing and sharing un

Short Text Language Identification for Under Resourced Languages

The paper presents a hierarchical naive Bayesian and lexicon based classifier for short text languag

Strategies for building wordnets for under-resourced languages: The case of African languages

The African Wordnet Project (AWN) aims at building wordnets for five African languages: Setswana, is

A Case Against Implicit Standards: Homophone Normalization in Machine Translation for Languages that use the Ge'ez Script

Homophone normalization, where characters that have the same sound in a writing script are mapped to

Translation-Based Dictionary Alignment for Under-Resourced Bantu Languages

Despite a large number of active speakers, most Bantu languages can be considered as under- or less-