Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

POS-tagging of Tunisian Dialect Using Standard Arabic Resources and Tools

Domain:

natural language processing

Record type:

paperdataset
Creator:
HamNasHabGal
Editor:
LabTraCenLab
Publisher:
CCSD
Host:avatar
International audience Developing natural language processing tools usually requires a large number of resources (lexica, annotated corpora, etc.), which often do not exist for less-resourced languages. One way to overcome the problem of lack of resources is to devote substantial efforts to build new ones from scratch. Another approach is to exploit existing resources of closely related languages. In this paper, we focus on developing a part-of-speech tagger for the Tunisian Arabic dialect (TUN), a low-resource language, by exploiting its close-ness to Modern Standard Arabic (MSA), which has many state-of-the-art resources and tools. Our system achieved an accuracy of 89% (∼20% absolute improvement over an MSA tagger baseline).

Visit

hal.science

Tasks

part of speech tagging

Languages

Arabic, Tunisian Spoken

Tags

[INFO.INFO-TT]Computer Science [cs]/Document and Text Processing

Licenses

info:eu-repo/semantics/OpenAccess