Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

A Survey on Multilingual Natural Language Processing: Data, Models, Evaluation, and Future Directions

Domain:

natural language processing

Record type:

paper
Creator:
Sri
Publisher:
Zenodo
Host:avatar
While natural language processing (NLP) has advanced for major languages, most of the world’s 7,000 languages remain underserved. I analyze 89 studies from 2016 to 2024, covering data-centric, model-centric, and evaluation methodologies for over 2,000 low-resource languages. I highlight initiatives like Glot500 (511 languages) and MasakhaNER (10 African languages). Despite progress in multilingual models, zero-shot cross-lingual transfer achieves only 60-70% of monolingual performance for typologically distant languages, and cultural appropriateness is under-evaluated. I propose culturally-aware evaluation, sustainable community partnerships, and parameter-efficient adaptation as priorities. Linguistic diversity is a resource for equitable NLP, requiring community collaboration.

Visit

doi.orgzenodo.org

Tags

Multilingual NLPLow-resource languagesMachine LearningNLP evaluation

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcodeCopyright (C) 2025 The Authors.http://rightsstatements.org/vocab/InC/1.0/