Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Multilingual and Cross-Lingual Citation Needed Detection on Wikipedia for Lower-Resource Languages

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
QuaRecBlaVra
Hôte:avatar
In automated fact-checking (AFC), check-worthiness detection identifies claims requiring verification based on domain-specific criteria. On Wikipedia, this task instantiates as Citation Needed Detection (CND), which flags claims lacking supporting citations. However, existing research has largely overlooked lower-resource languages, and recent AFC pipelines rely on large language models (LLMs), which are inaccessible to low-resource organizations. We introduce MCN, a multilingual CND corpus spanning 18 languages across three resource levels, on which we conduct an extensive study of small decoder-based language models (SLMs). Our experiments show that SLMs fine-tuned with an encoder-style objective substantially outperform prompted LLMs across languages. We further present one of the first studies on cross-lingual CND, demonstrating that SLMs fine-tuned solely on English claims surpass LLMs, even with little to no target-language adaptation. Our findings have important implications for lower-resource Wikipedia communities and suggest that compact, task-specific models are preferable to LLMs for CND. We release all data and code at github.com

Visit

arxiv.org

Tasks

text classification

Tags

Computation and Language

Similaires

Cross-lingual metaphor detection for low-resource languagesCross-Lingual Transfer Robustness to Lower-Resource Languages on Adversarial DatasetsCross-lingual transfer of multilingual models on low resource African LanguagesMultilingual Intermediate-Task Training for Zero-Shot Cross-Lingual Transfer on Low-Resource Languageskaranwxliaa/Cross-lingual-transfer-of-multilingual-models-on-low-resource-African-LanguagesCross-lingual Transfer Effects on Euphemism Detection in Low-Resource Languages

Cross-lingual metaphor detection for low-resource languages

State-of-the-art metaphor detection (MD) models achieve human-like performance for English data, whi

Cross-Lingual Transfer Robustness to Lower-Resource Languages on Adversarial Datasets

Multilingual Language Models (MLLMs) exhibit robust cross-lingual transfer capabilities, or the abil

Cross-lingual transfer of multilingual models on low resource African Languages

Large multilingual models have significantly advanced natural language processing (NLP) research. Ho

Multilingual Intermediate-Task Training for Zero-Shot Cross-Lingual Transfer on Low-Resource Languages

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

karanwxliaa/Cross-lingual-transfer-of-multilingual-models-on-low-resource-African-Languages

A comparison of Cross-lingual transfer of Transformer & Neural based: Multilingual & Monolingual mod

Cross-lingual Transfer Effects on Euphemism Detection in Low-Resource Languages

Euphemisms are culturally variable and often ambiguous, posing challenges for language models, espec