Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Scaling Frame Analysis to Genuinely Low-Resource Languages: A Case Study in Swahili and Tamil

Domaine:

natural language processing

Type de record:

datasetpaper
Créateur:
DepMarRC Dep
Éditeur:
Tec
Hôte:
Recent advances in multilingual news framing analysis have shown promise for low-resource settings through code-switching techniques. However, existing evaluations have focused on relatively high-resource languages like German, Turkish, and Arabic. This paper extends this line of research to genuinely low-resource languages—Swahili and Tamil—that face significant representation gaps in existing multilingual models. We introduce new annotated datasets for gun violence framing in these languages and systematically evaluate the code-switching approach under extreme low-resource conditions. Our results show that while code-switching provides consistent improvements over zero-shot transfer, the absolute performance gap between high-resource and genuinely low-resource languages remains substantial (15-20% F1-macro). We identify linguistic distance and morphological complexity as key challenges and propose adaptations to the code-switching method that yield 7% average improvement. Our work provides the first comprehensive analysis of cross-lingual frame detection in truly low-resource scenarios and establishes benchmarks for future research.

Visit

doi.org

Tasks

code switching

Languages

Swahili

Similaires

Introducing Syllable Tokenization for Low-resource Languages: A Case Study with SwahiliScaling Model Size and Zero-Shot Transfer to Low-Resource Languages in XTREMEZero-shot OCR Accuracy of Low-Resourced Languages: A Comparative Analysis on Sinhala and TamilScaling Pretrained Models and Intermediate-Task Training for Low-Resource Languages in XTREMEEnhancing Conversational AI for Low-Resource Languages: A Case Study on SomaliEnhancing Pos Tagging For Low-Resource Languages: A Case Study On Dholuo

Introducing Syllable Tokenization for Low-resource Languages: A Case Study with Swahili

Many attempts have been made in multilingual NLP to ensure that pre-trained language models, such as

Scaling Model Size and Zero-Shot Transfer to Low-Resource Languages in XTREME

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Zero-shot OCR Accuracy of Low-Resourced Languages: A Comparative Analysis on Sinhala and Tamil

Solving the problem of Optical Character Recognition (OCR) on printed text for Latin and its derivat

Scaling Pretrained Models and Intermediate-Task Training for Low-Resource Languages in XTREME

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Enhancing Conversational AI for Low-Resource Languages: A Case Study on Somali

Conversational AI has made huge strides in understanding and generating human language. However, the

Enhancing Pos Tagging For Low-Resource Languages: A Case Study On Dholuo