Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Impact of Pre-training Corpus Size on Cross-lingual Entity Recognition in Low-resource Languages

Domain:

natural language processing
Creator:
Ass
Publisher:
Zenodo
Host:avatar
Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to identify and classify named entities, making it particularly useful for low-resource languages. We show that the data-based cross-lingual transfer method is an effective technique for crosslingual NER and can outperform multilingual language models for low-resource languages. This paper introduces two key enhancements to the annotation projection step in cross-lingual NER for low-resource languages. First, we explore refining word alignments using back-translation to improve accuracy. Second, we pres Research goal: To what extent does increasing the pre-training corpus size of multilingual transformers like Bloom improve cross-lingual entity recognition accuracy on low-resource languages compared to XLM-R? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 7.5/10. This report was generated autonomously by Assignee Research, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 7.5/10.

Visit

doi.orgzenodo.org

Tasks

named entity recognitioninformation extraction

Tags

extentincreasingpre-trainingcorpussizemultilingualtransformerslike

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Impact of Parallel Corpus Size on Cross-Lingual Word Embedding Quality in Low-Resource LanguagesImpact of Intermediate Task Training Data Size on Zero-Shot Cross-Lingual Transfer Accuracy in Low-Resource LanguagesNamed Entity Recognition in Low-resource Languages using Cross-lingual distributional word representationMultimodal Teacher-Student Learning for Cross-Lingual Entity Recognition in Low-Resource LanguagesImpact of Intermediate-Task Training on Cross-Lingual Transfer Robustness in Low-Resource LanguagesMeta-Pretraining for Zero-Shot Cross-Lingual Named Entity Recognition in Low-Resource Philippine Languages

Impact of Parallel Corpus Size on Cross-Lingual Word Embedding Quality in Low-Resource Languages

We propose a new approach for learning contextualised cross-lingual word embeddings based on a small

Impact of Intermediate Task Training Data Size on Zero-Shot Cross-Lingual Transfer Accuracy in Low-Resource Languages

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Named Entity Recognition in Low-resource Languages using Cross-lingual distributional word representation

Named Entity Recognition (NER) is a fundamental task in many NLP applications that seek to identify

Multimodal Teacher-Student Learning for Cross-Lingual Entity Recognition in Low-Resource Languages

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident

Impact of Intermediate-Task Training on Cross-Lingual Transfer Robustness in Low-Resource Languages

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Meta-Pretraining for Zero-Shot Cross-Lingual Named Entity Recognition in Low-Resource Philippine Languages

Named-entity recognition (NER) in low-resource languages is usually tackled by finetuning very large