This paper presents a classification-based Vi-sual Question Answering (VQA) system for theHausa language, integrating Large LanguageModels (LLMs) and vision transformers. Byfine-tuning LLMs on monolingual Hausa textand fusing their representations with those ofstate-of-the-art vision encoders, our system pre-dicts answers from a fixed vocabulary. Exper-iments conducted on the HaVQA dataset, un-der offline text–image augmentation regimes,tailored to the specificity of Hausa as a low-resource language, show that this augmentationstrategy yields the best performance over thebaseline, achieving 35.85% accuracy, 35.89%WuPalmer similarity, and 15.32% F1-score.