Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Artificially Fluent: Swahili AI Performance Benchmarks Between English-Trained and Natively-Trained Datasets

Domain:

natural language processing

Record type:

paperdataset
Creator:
JafSay
Host:avatar
As large language models (LLMs) expand multilingual capabilities, questions remain about the equity of their performance across languages. While many communities stand to benefit from AI systems, the dominance of English in training data risks disadvantaging non-English speakers. To test the hypothesis that such data disparities may affect model performance, this study compares two monolingual BERT models: one trained and tested entirely on Swahili data, and another on comparable English news data. To simulate how multilingual LLMs process non-English queries through internal translation and abstraction, we translated the Swahili news data into English and evaluated it using the English-trained model. This approach tests the hypothesis by evaluating whether translating Swahili inputs for evaluation on an English model yields better or worse performance compared to training and testing a model entirely in Swahili, thus isolating the effect of language consistency versus cross-lingual abstraction. The results prove that, despite high-quality translation, the native Swahili-trained model performed better than the Swahili-to-English translated model, producing nearly four times fewer errors: 0.36% vs. 1.47% respectively. This gap suggests that translation alone does not bridge representational differences between languages and that models trained in one language may struggle to accurately interpret translated inputs due to imperfect internal knowledge representation, suggesting that native-language training remains important for reliable outcomes. In educational and informational contexts, even small performance gaps may compound inequality. Future research should focus on addressing broader dataset development for underrepresented languages and renewed attention to multilingual model evaluation, ensuring the reinforcing effect of global AI deployment on existing digital divides is reduced. 13 Pages, 3 Figures

Visit

arxiv.org

Languages

Swahili

Tags

Computation and LanguageComputers and Society

Similar

MehranLM-Tokenizer: A Natively-Trained Tokenizer for Sindhi Language ModellingPerformance of trained models.Scaling Performance of Models Trained on Artificially Code-Switched Data for Unseen Low-Resource Language PairsDisfluent-to-Fluent Tunisian Dialect Speech Translation with Fine-Tuning Pre-trained Language ModelsIndependent testing of ANN model trained on encoded datasets.Performance comparison of dense retrieval models trained on WebFAQ versus Wikipedia-based datasets for low-resource language

MehranLM-Tokenizer: A Natively-Trained Tokenizer for Sindhi Language Modelling

Sindhi, an Indo-Aryan language spoken by tens of millions of people, remains severely underrepresent

Performance of trained models.

Ensuring complete utilization of maternal continuum of care is essential for reducing matern

Scaling Performance of Models Trained on Artificially Code-Switched Data for Unseen Low-Resource Language Pairs

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Disfluent-to-Fluent Tunisian Dialect Speech Translation with Fine-Tuning Pre-trained Language Models

Independent testing of ANN model trained on encoded datasets.

A) Applying an ANN model trained on the Encoded-Muleba-GA dataset to estimate the parity status o

Performance comparison of dense retrieval models trained on WebFAQ versus Wikipedia-based datasets for low-resource language

We present WebFAQ, a large-scale collection of open-domain question answering datasets derived from