Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

A Bambara Tonalization System for Word Sense Disambiguation Using Differential Coding, Segmentation and Edit Operation Filtering

Domain:

natural language processing

Record type:

papersoftwaremodel
Creator:
LiuNou
Editor:
Équ
Publisher:
CCSD
Host:avatar
International audience In many languages such as Bambara or Arabic, tone markers (diacritics) may be written but are actually often omitted. NLP applications are confronted to ambiguities and subsequent difficulties when processing texts. To circumvent this problem , tonalization may be used, as a word sense disambiguation task, relying on context to add diacritics that partially disam-biguate words as well as senses. In this paper , we describe our implementation of a Bambara tonalizer that adds tone markers using machine learning (CRFs). To make our tool efficient, we used differential coding , word segmentation and edit operation filtering. We describe our approach that allows tractable machine learning and improves accuracy: our model may be learned within minutes on a 358K-word corpus and reaches 92.3% accuracy.

Visit

hal.science

Tasks

diacritic restorationtext normalization

Languages

BamanankanLame

Tags

ACM: I.: Computing Methodologies/I.2: ARTIFICIAL INTELLIGENCE/I.2.7: Natural Language Processing[INFO.INFO-LG]Computer Science [cs]/Machine Learning [cs.LG][INFO.INFO-CL]Computer Science [cs]/Computation and Language [cs.CL]

Licenses

info:eu-repo/semantics/OpenAccess