Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Investigating Bias in Bulgarian in the Context of Large Language Models

Domain:

natural language processing

Record type:

datasetpaper
Creator:
Ive
Publisher:
Spr
Host:
Abstract This paper investigates the detection and annotation of bias in Bulgarian with a view to large language models (LLMs). We present a Bias Annotated Dataset in Bulgarian comprising 3,177 sentences extracted from Bulgarian Wikipedia, manually annotated by two native speakers on five bias types (gender, religion, race and ethnicity, physical appearance, disability, and an additional ’Other’ category) using an ordinal scale from 0 to 5. We evaluate inter-annotator agreement both between the two human annotators and between humans and LLMs (Gemma 4, BgGPT-Gemma-3 with English and Bulgarian prompting, and Qwen 3) using weighted Cohen’s κ, Krippendorff’s α, and intraclass correlation coefficient (ICC). The results reveal high agreement between human annotators on the binary detection task (biased / non-biased classification) with mismatches of only 4.9%, but substantially lower agreement on intensity scoring, consistent with the inherently subjective nature of bias perception. LLMs show significant divergence from human judgements, with much higher mismatch rates of 24–34% and uniformly negative agreement metrics, indicating that LLMs flag largely different sentences as biased and assign systematically different (lower) intensity scores than human annotators. With respect to the BgGPT- Gemma-3 models specifically focused on Bulgarian, the results show that prompting in Bulgarian rather than English improves LLM alignment with human annotation, highlighting the importance of target-language prompting for low-resource and culturally specific bias detection. The findings contribute to the understanding of bias in Bulgarian and low-resource languages more broadly, and underscore the need for culturally grounded lexical resources and annotation methodologies.

Visit

doi.org

Tasks

text classification

Licenses

https://creativecommons.org/licenses/by/4.0/

Similar

Sociolinguistic Bias and Language Inequality in Large Language ModelsAnalyzing In-Context Language Learning in Long-Context Large Language ModelsInvestigating Cultural Alignment of Large Language ModelsEvaluating Racial Bias in Large Language Models: The Necessity for "SMOKY"Context-Aware Large Language Models for Multilingual UnderstandingCross-Lingual Bias in Large Language Models: A Comparative Analysis of English and Swahili

Sociolinguistic Bias and Language Inequality in Large Language Models

Large language models (LLMs) are increasingly deployed across multilingual applications, yet persist

Analyzing In-Context Language Learning in Long-Context Large Language Models

We evaluated the in-context learning capabilities of long-context large language models for machine

Investigating Cultural Alignment of Large Language Models

The intricate relationship between language and culture has long been a subject of exploration withi

Evaluating Racial Bias in Large Language Models: The Necessity for "SMOKY"

This paper evaluates the understanding and biases of large language models (LLMs) regarding

Context-Aware Large Language Models for Multilingual Understanding

Multilingual large language models (LLMs) have demonstrated strong performance in cross-lingual task

Cross-Lingual Bias in Large Language Models: A Comparative Analysis of English and Swahili