Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

KZ-SafetyPrompts: A Kazakh Safety Evaluation Prompt Dataset for Large Language Models

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
zagIbrMurZha
Hôte:avatar
Kazakh is underrepresented in resources for evaluating the safety behavior of large language models. We present KZ-SafetyPrompts, a Kazakh prompt dataset for safety evaluation across eleven categories covering common risk areas such as self-harm, violence, child exploitation, sexual content, racist content, radicalization, and regulated goods or illegal activities. The dataset contains 5,717 prompts written natively in Kazakh (Cyrillic), organized by category, with English translations for cross-lingual analysis. Prompts resemble realistic user queries, often in a teen or child style, and are phrased as intent prompts without procedural instructions. We document the writing protocol, labeling procedures (including borderline-case decision rules), and quality-control steps (schema standardization, completeness checks, and deduplication). We also align the categories with widely used safety taxonomies to support integration with existing evaluation pipelines. Baseline results with GPT-4o show an overall refusal rate of 28.2%, varying from 5.5% to 53.8% across categories, indicating that Kazakh prompts expose category-specific safety gaps not captured by English-only evaluation. Accepted at the SIGUL2026 Workshop co-located with LREC2026

Visit

arxiv.org

Tasks

text classification

Tags

Computation and Language

Similaires

Prompt reinforcing for long-term planning of large language modelsEthiopicEmotion: Multi-label Emotion Dataset with Large Language Models EvaluationMultilingual Prompt Engineering in Large Language Models: A Survey Across NLP TasksPrompt engineering on large language models (LLMs) in low-resourced language settingMulti-lingual Functional Evaluation for Large Language ModelsEvaluation Mirage: A Layered Evaluation of Large Language Models and Language Identification for African NLP

Prompt reinforcing for long-term planning of large language models

Large language models (LLMs) have achieved remarkable success in a wide range of natural language pr

EthiopicEmotion: Multi-label Emotion Dataset with Large Language Models Evaluation

Large Language Models (LLMs) show promising learning and reasoning abilities. Compared to other NLP

Multilingual Prompt Engineering in Large Language Models: A Survey Across NLP Tasks

Large language models (LLMs) have demonstrated impressive performance across a wide range of Natural

Prompt engineering on large language models (LLMs) in low-resourced language setting

The attached dataset has all the information in regard to Large Language Models (LLMs)

Multi-lingual Functional Evaluation for Large Language Models

Multi-lingual competence in large language models is often evaluated via static data benchmarks such

Evaluation Mirage: A Layered Evaluation of Large Language Models and Language Identification for African NLP

David Ifeoluwa Adelani (Supervisor) As Large Language Models (LLMs) are increasingly deployed in glo