Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

BRIDGING GAPS IN LOW-RESOURCE LANGUAGES: A MACHINE LEARNING APPROACH TO PRONOMINAL ANAPHORA RESOLUTION

Domaine:

natural language processing

Type de record:

paper
Créateur:
Kal
Éditeur:
Sou
Hôte:
This study proposes the Kazakh Coreference Adaptation (KCA) model, a hybrid framework for resolving pronominal anaphora in Kazakh, a morphologically rich and low-resource language for which existing coreference systems and large language models (LLMs) remain unreliable. The study aims to evaluate whether integrating explicit linguistic constraints with supervised machine learning can outperform large language models in this setting. The novelty of the approach lies in combining deterministic morphological filtering with probabilistic ranking trained on a newly annotated subset of the Kazakh National Corpus. The model integrates rule-based constraints-covering case, number, person agreement, and syntactic accessibility-with a RandomForest classifier trained on 4,200 sentences across five genres. The linguistic stage eliminates grammatically implausible candidates, while the machine-learning component ranks the remaining antecedents using morphological, syntactic, discourse, and semantic features. Experimental results show that KCA achieves 79% accuracy on long-context anaphora cases, exceeding GPT-4o (70%) and o1-mini (73%). The hybrid architecture demonstrates clear advantages in resolving long-distance and morphologically complex structures. These findings highlight the importance of morphology-aware modeling in agglutinative languages and provide a reproducible dataset and methodological framework applicable to other low-resource settings.

Visit

doi.org

Tasks

coreference resolution

Similaires

Neural pronominal anaphora resolution for Amharic text using multiple embeddingsAmharic Anaphora Resolution Using Knowledge-Poor ApproachText Detoxification in isiXhosa and Yorùbá: A Cross-Lingual Machine Learning Approach for Low-Resource African LanguagesInstituteforDiseaseModeling/Bridging-the-Gap-Low-Resource-African-LanguagesYaregalAssabie/Amharic-Anaphora-ResolutionAdapting Machine Learning Techniques for Low-Resource Settings in Developing Countries: A Multidisciplinary Approach

Neural pronominal anaphora resolution for Amharic text using multiple embeddings

Amharic Anaphora Resolution Using Knowledge-Poor Approach

Text Detoxification in isiXhosa and Yorùbá: A Cross-Lingual Machine Learning Approach for Low-Resource African Languages

Toxic language is one of the major barrier to safe online participation, yet robust mitigation tools

InstituteforDiseaseModeling/Bridging-the-Gap-Low-Resource-African-Languages

Code and data corresponding to paper Bridging the Gap: Enhancing LLM Performance for Low-Resource Af

YaregalAssabie/Amharic-Anaphora-Resolution

# Amharic-Anaphora-Resolution This is a readme file.

Adapting Machine Learning Techniques for Low-Resource Settings in Developing Countries: A Multidisciplinary Approach

Developing countries face unique challenges in harnessing the power of machine learning (ML) due to