Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

BERT-based Models vs. Large Language Models for Low-Resource Named Entity Recognition: A Comparative Study on Marathi

Domaine:

natural language processing

Type de record:

paper
Créateur:
IngGhoGonHar
Éditeur:
arXiv
Hôte:avatar
Named Entity Recognition (NER) for low-resource languages such as Marathi remains a challenging task due to limited annotated resources and linguistic complexity. Although recent Large Language Models (LLMs) have demonstrated strong performance across a wide range of natural language processing tasks, their effectiveness for language-specific NER in low-resource settings remains uncertain. In this study, we fine-tune MahaBERT-v2 on different variants of the MahaNER dataset and systematically compare the performance of these models with an existing MahaNER baseline and prominent general-purpose LLMs, including Gemini, LLaMA-3.3-70B, and Gemma models. All models are evaluated on a Marathi NER test dataset using standard metrics of precision, recall, and F1-score. The experimental results show that the fine-tuned MahaBERT-based models consistently outperform both the baseline and all evaluated LLMs, with the fine-tuned models achieving F1-scores ranging from 0.88 to 0.91, surpassing the existing MahaNER model (0.8843) and significantly exceeding the performance of LLM-based approaches, whose F1-scores range from 0.57 to 0.69. These findings demonstrate that task-specific, language-focused models trained on domain-relevant data remain more effective than general-purpose LLMs for Marathi NER, highlighting the continued importance of specialized architectures for low-resource language processing.

Visit

doi.org

Tasks

information extractionnamed entity recognition

Tags

Computation and Language (cs.CL)Machine Learning (cs.LG)FOS: Computer and information sciences

Licenses

arXiv.org perpetual, non-exclusive licensehttp://arxiv.org/licenses/nonexclusive-distrib/1.0/

Similaires

Mono vs Multilingual BERT: A Case Study in Hindi and Marathi Named Entity RecognitionRetrieveAll: A Multilingual Named Entity Recognition Framework with Large Language ModelsNamed Entity Recognition in Arabic Mental Health Using Large Language ModelsOn the Strength of Character Language Models for Multilingual Named Entity RecognitionA Hybrid Method for Low-Resource Named Entity RecognitionANEA: Distant Supervision for Low-Resource Named Entity Recognition

Mono vs Multilingual BERT: A Case Study in Hindi and Marathi Named Entity Recognition

Named entity recognition (NER) is the process of recognising and classifying important information (

RetrieveAll: A Multilingual Named Entity Recognition Framework with Large Language Models

The rise of large language models has led to significant performance breakthroughs in named entity r

Named Entity Recognition in Arabic Mental Health Using Large Language Models

Named Entity Recognition (NER) in Arabic mental-health text is constrained by the scarcity of gold a

On the Strength of Character Language Models for Multilingual Named Entity Recognition

Character-level patterns have been widely used as features in English Named Entity Recognition (NER)

A Hybrid Method for Low-Resource Named Entity Recognition

Named Entity Recognition (NER) is a critical component of Natural Language Processing with diverse a

ANEA: Distant Supervision for Low-Resource Named Entity Recognition

Distant supervision allows obtaining labeled training corpora for low-resource settings where only l