Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

A Hybrid Statistical and Rule-based Approach to Extremely Low-resource Machine Transliteration

Domaine:

natural language processing

Type de record:

paper
Créateur:
Pat
Éditeur:
Ass
Hôte:
Machine transliteration work has focused primarily on languages with large volumes of parallel corpus, and between language pairs whose orthographies are very different. In contrast, a large proportion of the world’s languages have vastly fewer resources and employ Roman-like alphabets often with large degrees of orthographic overlap with high-resource languages. We propose that machine transliteration between languages with few training examples can be accomplished by a noisy-channel-like statistical model captured in a human editable format with practical rule-based capabilities built-in. This hybrid approach allows users to take advantage of an algorithm to find and apply common transformations in context while providing rigorous control over the output. Effectiveness is evaluated on the Bible names translation matrix dataset of Wu et al. (2018), covering 591 languages that involve 590 names on average per language pair. Our approach slightly exceeds past results and explores several features targeted at benefiting the extremely low-resource language domain.

Visit

doi.org

Tasks

text normalization

Licenses

https://www.acm.org/publications/policies/copyright_policy#Background

Similaires

Enhancing Cross-Lingual Transfer through Reversible Transliteration: A Huffman-Based Approach for Low-Resource LanguagesA Rule-based Approach to English-Okun Prepositional Phrase Machine TranslationA Hybrid Approach to Contextual Information Extraction in Low-Resource IgboSinhala Transliteration: A Comparative Analysis Between Rule-based and Seq2Seq ApproachesHybrid Multilingual Models with Vocabulary Augmentation and Script Transliteration for Low-Resource POS TaggingCharSpan: Utilizing Lexical Similarity to Enable Zero-Shot Machine Translation for Extremely Low-resource Languages

Enhancing Cross-Lingual Transfer through Reversible Transliteration: A Huffman-Based Approach for Low-Resource Languages

As large language models (LLMs) are trained on increasingly diverse and extensive multilingual corpo

A Rule-based Approach to English-Okun Prepositional Phrase Machine Translation

Okun culture is gradually going into extinction because the language is greatly dominated by English

A Hybrid Approach to Contextual Information Extraction in Low-Resource Igbo

Extracting contextual information from low-resource languages such as Igbo remains a significant cha

Sinhala Transliteration: A Comparative Analysis Between Rule-based and Seq2Seq Approaches

Due to reasons of convenience and lack of tech literacy, transliteration (i.e., Romanizing native sc

Hybrid Multilingual Models with Vocabulary Augmentation and Script Transliteration for Low-Resource POS Tagging

Pretrained multilingual language models have become a common tool in transferring NLP capabilities t

CharSpan: Utilizing Lexical Similarity to Enable Zero-Shot Machine Translation for Extremely Low-resource Languages

We address the task of machine translation (MT) from extremely low-resource language (ELRL) to Engli