Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Domain Adaptation in Sequence Labelling: A Case Study for Two South African Languages

Domaine:

natural language processing

Type de record:

paper
Créateur:
TanRoald Eiselen
Éditeur:
Lin
Hôte:
In this paper, we investigate domain adaptation for Part-of-speech (POS) tagging of two under-resourced South African languages, isiZulu and Sesotho sa Leboa, by studying its effect on the POS tagging results and how to possibly predict what quality can be expected when applying an existing POS tagger to a new domain. We carry out systematic experiments across six domains (governmental texts, exam texts for grade 12 South African learners, magazines, newspapers, novels, and PhD theses) to determine how POS tagger accuracy deteriorates when switching between domains. To mitigate this quality deterioration, three different domain adaptation strategies are tested to determine the most relevant approach in highly under-resourced scenarios. The results of these experiments show that adding even relatively small amounts of annotated data from a target domain delivers the highest accuracy on the target domain compared to other domain adaptation methods. To determine the underlying causes of the accuracy deterioration, a forward stepwise linear regression modelling experiment shows that a combination of lexical and syntactic divergence can account for a significant amount of the deterioration, and are good predictors of the expected deterioration when applying POS tagging to a new domain.

Visit

doi.org

Tasks

part of speech taggingtransfer learning

Languages

Sotho, NorthernSotho, SouthernZulu

Licenses

https://creativecommons.org/licenses/by/4.0

Similaires

AdaSL: An Unsupervised Domain Adaptation framework for Arabic multi-dialectal Sequence LabelingGovernment Domain Named Entity Recognition for South African LanguagesGoverning climate adaptation innovation in Africa: A South African case studyDomain Adaptation Effects in Cross-Lingual NER for Low-Resource LanguagesLarge Language Models Adaptation for Low-resource Languages: The Case for African LanguagesClassifier identification in Ancient Egyptian as a low-resource sequence-labelling task

AdaSL: An Unsupervised Domain Adaptation framework for Arabic multi-dialectal Sequence Labeling

Government Domain Named Entity Recognition for South African Languages

This paper describes the named entity language resources developed as part of a development project for the South African languages. The development efforts focused on creating protocols and annotated data sets with at least 15,000 annotated named entity tokens for

Governing climate adaptation innovation in Africa: A South African case study

Despite contributing little to global warming, Africa continues to be adversely impacted by climate

Domain Adaptation Effects in Cross-Lingual NER for Low-Resource Languages

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident

Large Language Models Adaptation for Low-resource Languages: The Case for African Languages

David Ifeoluwa Adelani (Supervisor) Despite remarkable advances in Large Language Models (LLMs), Afr

Classifier identification in Ancient Egyptian as a low-resource sequence-labelling task

The complex Ancient Egyptian (AE) writing system was characterised by widespread use of graphemic cl