Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

L3Cube-MahaPOS: A Marathi Part-of-Speech Tagging Dataset and BERT Models

Domaine:

natural language processing

Type de record:

paperdatasetmodel
Créateur:
IngGhoGonHar
Hôte:avatar
Part-of-Speech (POS) tagging is a foundational NLP task underpinning machine translation, information extraction, and syntactic parsing. Despite Marathi being spoken by over 83 million people and ranking among the top twenty most spoken languages worldwide, it remains severely under-resourced in annotated corpora and standardised evaluation benchmarks. Marathi presents unique challenges for computational modelling owing to its rich morphology, relatively free word order, lack of capitalisation conventions, and pervasive code-mixing with Hindi and English. We introduce L3Cube-MahaPOS, a gold-standard POS tagging dataset for Marathi comprising 32,354 manually annotated sentences drawn from news text. Annotation was performed entirely manually by a team of Marathi-proficient annotators following a 16-tag Universal Dependencies-aligned scheme. A structured preprocessing pipeline covering Unicode normalisation, Devanagari-aware tokenisation, and noise filtering ensures label consistency across all splits. We benchmark the dataset across six model families spanning HMM, CRF, BiLSTM, BiLSTM+CharCNN, MuRIL, and the Marathi-specific transformer MahaBERT-v2. The best system achieves 88.67\% token-level accuracy and a macro-F1 of 81.67% over 15 evaluated tag classes. We release the dataset, annotation guidelines, and trained model checkpoints to foster further research in Marathi NLP.

Visit

arxiv.org

Tasks

part of speech tagging

Tags

Computation and LanguageMachine Learning

Similaires

L3Cube-MahaSTS: A Marathi Sentence Similarity Dataset and ModelsEnhanced Part-of-Speech Tagging Resources and Models for Tigrinyanaftalindeapo/Part-of-speech-tagging-with-hidden-Markov-modelsMahaParaphrase: A Marathi Paraphrase Detection Corpus and BERT-based ModelsSetswana Part of Speech TaggingMouhamedkhlifi/Part-of-speech-tagging

L3Cube-MahaSTS: A Marathi Sentence Similarity Dataset and Models

We present MahaSTS, a human-annotated Sentence Textual Similarity (STS) dataset for Marathi, along w

Enhanced Part-of-Speech Tagging Resources and Models for Tigrinya

naftalindeapo/Part-of-speech-tagging-with-hidden-Markov-models

This project aims to develop a part-of-speech tagger for Afrikaans, a South African language, using

MahaParaphrase: A Marathi Paraphrase Detection Corpus and BERT-based Models

Paraphrases are a vital tool to assist language understanding tasks such as question answering, styl

Setswana Part of Speech Tagging

Mouhamedkhlifi/Part-of-speech-tagging

Training xlmroberta model on 20 typologically diverse african languages to classify 14 parts of spee