Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Low-Resource Indonesian Health Conversations

Domaine:

natural language processinghealthcare

Type de record:

dataset
Créateur:
SuwNur
Éditeur:
Zenodo
Hôte:avatar
The dataset consists of 5,000 question–answer (QA) pairs collected and annotated for low-resource Indonesian health conversations in a child education context. Each QA pair represents a natural interaction between elementary school students and a conversational health system, covering everyday health topics such as hygiene, nutrition, exercise, and disease prevention. The queries are characterized by short sentence structures, simple vocabulary, and direct questioning patterns, reflecting the linguistic and cognitive profiles of children. The corresponding responses are designed to be concise, informative, and pedagogically appropriate, providing clear explanations that support understanding rather than merely delivering factual answers. All data instances are annotated using a hierarchical intent schema, consisting of coarse-grained categories and fine-grained sub-intents. This structure enables the modeling of semantic relationships between health concepts and supports more accurate intent classification in ambiguous or overlapping cases. Despite its moderate size, the dataset reflects a realistic low-resource scenario, where annotated conversational data in Indonesian remains limited. At the same time, the scale of 5,000 QA pairs provides sufficient diversity and coverage to support robust training and evaluation of transformer-based conversational models. Overall, the dataset captures both the simplicity of child-oriented language and the semantic complexity of health-related queries, making it suitable for research on hierarchical intent modeling, context-aware dialogue systems, and explainable NLP in educational settings.

Visit

doi.orgzenodo.org

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

Low-Resource Corpus Indonesian Local LanguageConstructing and Expanding Low-Resource and Underrepresented Parallel Datasets for Indonesian Local LanguagesNusaMT-7B: Machine Translation for Low-Resource Indonesian Languages with Large Language ModelsMultilingual Idioms in Sentences and Conversations Across High-, Medium-, and Low-Resource LanguagesSLANG-GraphRAG: Multi-Layered Retrieval with Domain-Specific Knowledge for Low Resource Social Media ConversationsHealth QA in Low-Resource African Languages

Low-Resource Corpus Indonesian Local Language

This study departs from the hypothesis that combining Neural Machine Translation (NMT) with the stem

Constructing and Expanding Low-Resource and Underrepresented Parallel Datasets for Indonesian Local Languages

In Indonesia, local languages play an integral role in the culture. However, the available Indonesia

NusaMT-7B: Machine Translation for Low-Resource Indonesian Languages with Large Language Models

Large Language Models (LLMs) have demonstrated exceptional promise in translation tasks for high-res

Multilingual Idioms in Sentences and Conversations Across High-, Medium-, and Low-Resource Languages

Idiomatic expressions pose a major challenge for multilingual NLP because their meanings shift betwe

SLANG-GraphRAG: Multi-Layered Retrieval with Domain-Specific Knowledge for Low Resource Social Media Conversations

Emotion classification on social media is especially difficult when texts include informal, cultural

Health QA in Low-Resource African Languages

Maternal, sexual and reproductive health QA pairs across four African languages