Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Impact of Pre-training Corpus Size on Cross-lingual Entity Recognition in Low-resource Languages

Domaine:

natural language processing
Créateur:
Ass
Éditeur:
Zenodo
Hôte:avatar
Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to identify and classify named entities, making it particularly useful for low-resource languages. We show that the data-based cross-lingual transfer method is an effective technique for crosslingual NER and can outperform multilingual language models for low-resource languages. This paper introduces two key enhancements to the annotation projection step in cross-lingual NER for low-resource languages. First, we explore refining word alignments using back-translation to improve accuracy. Second, we pres Research goal: To what extent does increasing the pre-training corpus size of multilingual transformers like Bloom improve cross-lingual entity recognition accuracy on low-resource languages compared to XLM-R? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 7.5/10. This report was generated autonomously by Assignee Research, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 7.5/10.

Visit

doi.orgzenodo.org

Tasks

named entity recognitioninformation extraction

Tags

extentincreasingpre-trainingcorpussizemultilingualtransformerslike

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

Impact of Parallel Corpus Size on Cross-Lingual Word Embedding Quality in Low-Resource LanguagesImpact of Intermediate Task Training Data Size on Zero-Shot Cross-Lingual Transfer Accuracy in Low-Resource LanguagesNamed Entity Recognition in Low-resource Languages using Cross-lingual distributional word representationMultimodal Teacher-Student Learning for Cross-Lingual Entity Recognition in Low-Resource LanguagesImpact of Intermediate-Task Training on Cross-Lingual Transfer Robustness in Low-Resource LanguagesMeta-Pretraining for Zero-Shot Cross-Lingual Named Entity Recognition in Low-Resource Philippine Languages

Impact of Parallel Corpus Size on Cross-Lingual Word Embedding Quality in Low-Resource Languages

We propose a new approach for learning contextualised cross-lingual word embeddings based on a small

Impact of Intermediate Task Training Data Size on Zero-Shot Cross-Lingual Transfer Accuracy in Low-Resource Languages

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Named Entity Recognition in Low-resource Languages using Cross-lingual distributional word representation

Named Entity Recognition (NER) is a fundamental task in many NLP applications that seek to identify

Multimodal Teacher-Student Learning for Cross-Lingual Entity Recognition in Low-Resource Languages

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident

Impact of Intermediate-Task Training on Cross-Lingual Transfer Robustness in Low-Resource Languages

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Meta-Pretraining for Zero-Shot Cross-Lingual Named Entity Recognition in Low-Resource Philippine Languages

Named-entity recognition (NER) in low-resource languages is usually tackled by finetuning very large