Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

DATASHI: A Parallel English-Tashlhiyt Corpus for Orthography Normalization and Low-Resource Language Processing

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
MonBao
Hôte:avatar
DATASHI is a new parallel English-Tashlhiyt corpus that fills a critical gap in computational resources for Amazigh languages. It contains 5,000 sentence pairs, including a 1,500-sentence subset with expert-standardized and non-standard user-generated versions, enabling systematic study of orthographic diversity and normalization. This dual design supports text-based NLP tasks - such as tokenization, translation, and normalization - and also serves as a foundation for read-speech data collection and multimodal alignment. Comprehensive evaluations with state-of-the-art Large Language Models (GPT-5, Claude-Sonnet-4.5, Gemini-2.5-Pro, Mistral, Qwen3-Max) show clear improvements from zero-shot to few-shot prompting, with Gemini-2.5-Pro achieving the lowest word and character-level error rates and exhibiting robust cross-lingual generalization. A fine-grained analysis of edit operations - deletions, substitutions, and insertions - across phonological classes (geminates, emphatics, uvulars, and pharyngeals) further highlights model-specific sensitivities to marked Tashlhiyt features and provides new diagnostic insights for low-resource Amazigh orthography normalization. This paper has been accepted for presentation at LREC 2026

Visit

arxiv.org

Tasks

text normalization

Languages

AmazighBerberTachelhit

Tags

Computation and Language

Similaires

TWIENG: A Multi-Domain Twi-English Parallel Corpus for Machine Translation of Twi, a Low-Resource African LanguageTwieng: A Multi-Domain Twi-English Parallel Corpus for Machine Translation of the Twi Language, A Low-Resource African LanguageA Low-Resource English–Hassaniya Parallel Corpus with Neural Machine Translation BenchmarksRIYE Audio Dataset: A Multidialectal Speech Corpus for Low-Resource Language ProcessingParaKab – Many Languages, One Kabyle: A Multilingual Parallel Corpus for a Low-Resource LanguageEthioMT: Parallel Corpus for Low-resource Ethiopian Languages

TWIENG: A Multi-Domain Twi-English Parallel Corpus for Machine Translation of Twi, a Low-Resource African Language

A Twi-English parallel corpus is certainly an important resource for Machine Translation of Twi (ISO

Twieng: A Multi-Domain Twi-English Parallel Corpus for Machine Translation of the Twi Language, A Low-Resource African Language

A Twi-English parallel corpus is certainly an important resource for Machine Translation of Twi (ISO

A Low-Resource English–Hassaniya Parallel Corpus with Neural Machine Translation Benchmarks

RIYE Audio Dataset: A Multidialectal Speech Corpus for Low-Resource Language Processing

This dataset consists of a curated collection of high-fidelity, field-recorded audio samples develop

ParaKab – Many Languages, One Kabyle: A Multilingual Parallel Corpus for a Low-Resource Language

Description of the Dataset This dataset consists of three parallel corpora involving the Kabyle lan

EthioMT: Parallel Corpus for Low-resource Ethiopian Languages

Recent research in natural language processing (NLP) has achieved impressive performance in tasks such as machine translation (MT), news classification, and question-answering in high-resource languages. However, the performance of MT leaves much to be desired for