Large language models (LLMs) are increasingly deployed across multilingual applications, yet persistent disparities in performance and fairness raise concerns about equitable access to generative AI. This paper investigates sociolinguistic bias and language inequality in LLMs across six languages: English, French, and Mandarin Chinese as high-resource languages, and Yorùbá, Swahili, and Haitian Creole as lower-resource languages. It also examines three English-based code-switching settings: English-Yorùbá, English-Swahili, and English-Haitian Creole.
Using a controlled evaluation framework comprising stereotype-eliciting prompts, neutral controls, factual question answering, creative generation, and code-switching prompts, the study measures stereotype replication, harmful generalization, fluency, factual accuracy, and reliability across repeated samples. The core empirical dataset contains 17,100 non-mitigated model completions generated from five instruction-tuned systems. Mitigation analyses were conducted as a paired evaluation on the same prompt framework and are reported separately.
The analysis combines human annotation, multilingual sentence embedding association tests, and statistical hypothesis testing using logistic regression and hierarchical linear models. Results show that LLMs replicate stereotypes across all evaluated languages, with higher bias incidence and weaker mitigation effects in lower-resource languages. Performance degradation is both quantitative, through lower mean fluency and accuracy, and qualitative, through mistranslations, hallucinations, language collapse, and greater output instability. Code-switching substantially increases error rates and output variance. Prompt engineering reduces measured bias but does not close cross-lingual gaps, suggesting that disparities are structurally tied to training data, alignment coverage, and evaluation regimes rather than user-level prompting alone.