Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Pretraining Corpus Domain and Hybrid Multilingual Model Performance in Low-Resource Universal Dependency Parsing

Domaine:

natural language processing

Type de record:

paper
Créateur:
Ass
Éditeur:
Zenodo
Hôte:avatar
Pretrained multilingual language models have become a common tool in transferring NLP capabilities to low-resource languages, often with adaptations. In this work, we study the performance, extensibility, and interaction of two such adaptations: vocabulary augmentation and script transliteration. Our evaluations on part-of-speech tagging, universal dependency parsing, and named entity recognition in nine diverse low-resource languages uphold the viability of these approaches while raising new questions around how to optimally adapt multilingual models to low-resource settings. Research goal: To what extent does the choice of pretraining corpus (domain-specific vs. general) influence the performance of hybrid multilingual models in universal dependency parsing for low-resource languages, as measured by labeled attachment score (LAS)? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 7.7/10. This report was generated autonomously by Assignee Research, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 7.7/10.

Visit

doi.orgzenodo.org

Tasks

dependency parsinginformation extractionnamed entity recognitionparsingpart of speech tagging

Tags

extentchoicepretrainingcorpusdomain-specificgeneralinfluenceperformance

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode