Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Multimodal Classification System for Hausa Using LLMs and Vision Transformers

Domaine:

natural language processing

Type de record:

paper
Créateur:
AliFat
Éditeur:
Dig
Hôte:
This paper presents a classification-based Vi-sual Question Answering (VQA) system for theHausa language, integrating Large LanguageModels (LLMs) and vision transformers. Byfine-tuning LLMs on monolingual Hausa textand fusing their representations with those ofstate-of-the-art vision encoders, our system pre-dicts answers from a fixed vocabulary. Exper-iments conducted on the HaVQA dataset, un-der offline text–image augmentation regimes,tailored to the specificity of Hausa as a low-resource language, show that this augmentationstrategy yields the best performance over thebaseline, achieving 35.85% accuracy, 35.89%WuPalmer similarity, and 15.32% F1-score.

Visit

doi.org

Tasks

question answering

Languages

Hausa

Licenses

https://creativecommons.org/licenses/by-sa/4.0