Logo Lanfrica

Dataset for "A Benchmark Suite for Indonesian Mental Health NER: Combining Real Counselling Audio Transcripts and Adapted Text Data"

Domaine:

natural language processinghealthcare

Type de record:

dataset
Créateur:
AmaLydRanTan
Éditeur:
Zenodo
Hôte:avatar
Description This dataset focuses on utilizing primary data in the form of original Indonesian psychological counseling audio transcripts as the main source. These data constitute the core contribution of the dataset, as they represent authentic conversations between clients and professionals, which have been very limited in Indonesian. To complement language coverage and variability, this dataset also incorporates secondary data in the form of public mental health texts adapted from social media. Overall, the corpus in this dataset comprises 23,854 utterances with a total of 2,422,402 tokens, of which 237,590 tokens are specifically annotated for the Named Entity Recognition (NER) task using the CoNLL-2003 format and the BIO labeling scheme. The annotation is designed specifically for the mental health domain, covering eight entity categories. The main strength of this dataset lies in the use of primary audio data transcribed using the Whisper model, enabling it to capture the characteristics of natural spoken language, including emotional expressions, irregular sentence structures, and real conversational context. Within a single utterance, clients may express multiple entities simultaneously, such as PROFESSIONAL ("doctor"), TRIGGER ("my job is quite hard"), SYMPTOM ("weak"), and IDENTIFICATION ("stress"), reflecting the complexity of real-world data. To ensure quality and privacy, this dataset applies a rigorous de-identification process to mask sensitive information and leverages Google Translate to enrich the data through adaptation of peer-support texts originally in English. With this approach, the dataset provides a rich and realistic resource while addressing the scarcity of annotated clinical dialogue data in Indonesian, a low-resource language. This dataset is designed to evaluate the robustness of several pretrained language models on mental health texts and to facilitate the development of automated information extraction systems for psychologists’ post-session documentation. It is introduced under the title A Benchmark Suite for Indonesian Mental Health NER: Combining Real Counseling Audio Transcripts and Adapted Text Data.   Annotation Methodology The annotation methodology in this study is designed in a structured and systematic manner to ensure consistency and accuracy in extracting psychological information from textual data and audio transcripts. The annotation process is carried out through the following main stages: Entity Schema DevelopmentThe entity schema is developed iteratively through literature review, preliminary corpus inspection, and consultation with domain experts in psychology. This process results in eight primary psychological entity categories: IDENTIFICATION, SYMPTOM, EMOTION, TRIGGER, TREATMENT, MEDICATION, PATIENT, and PROFESSIONAL. Operational definitions are established for each category to ensure consistent interpretation throughout the annotation process. Data Preparation and De-identificationAs the basis for annotation (gold standard), this study uses transcripts that have undergone limited normalization (ASR-Clean). Before being distributed to annotators, the data is subjected to a strict de-identification process to protect participant privacy and obscure any sensitive information that could reveal individual identities. Tokenization and Tagging FormatEach utterance in the corpus is processed through tokenization to segment the text into discrete token units. Each token is then labeled using the BIO (Beginning, Inside, Outside) tagging scheme. The annotated data is stored in the CoNLL-2003 format, where each line represents a token along with its entity label, and utterance boundaries are separated by blank lines. This format is chosen due to its compatibility with various machine learning frameworks for Named Entity Recognition tasks. Minimal-Span PrincipleThe annotation guidelines strictly apply the minimal-span principle. Annotators are instructed to label only the text segments that directly represent the target entity, without including non-essential surrounding words. This approach aims to improve annotation precision and reduce inter-annotator variability. Ambiguity Handling and CalibrationGiven the complexity of narratives in the psychological domain, particular attention is paid to potential overlaps between entity categories, such as SYMPTOM, EMOTION, and TRIGGER, as well as between TREATMENT and MEDICATION. In such cases, annotation decisions are based on the most dominant meaning within the context of the utterance. If ambiguity arises, annotators assign a special marker for later discussion during the calibration (adjudication) phase to ensure consistent annotation decisions.   Label Definitions The entity schema is designed to capture key information commonly found in mental health narratives and counseling sessions. This dataset defines eight primary entity categories (labels) as follows: IDENTIFICATIONDescription: Psychological conditions or diagnoses.Examples: stress, depression, burnout. SYMPTOMDescription: Mental, physical, or emotional symptoms experienced by the subject.Examples: fatigue, difficulty sleeping, loss of appetite. EMOTIONDescription: Emotions explicitly expressed by the subject.Examples: sadness, anxiety, fear. TRIGGERDescription: Factors that trigger psychological distress.Examples: heavy workload, family problems, academic pressure. TREATMENTDescription: Therapeutic actions or interventions undertaken.Examples: counseling, therapy, relaxation exercises. MEDICATIONDescription: Medications or supplements used in the treatment process.Examples: sleeping pills, antidepressants, calming supplements. PATIENTDescription: Pronouns or references to the subject or patient being discussed.Examples: I, my child, this patient. PROFESSIONALDescription: Mental health professionals or related institutions.Examples: doctor, psychologist, hospital.   Data Partitioning For the purpose of training and evaluating Named Entity Recognition (NER) models, the fully annotated corpus is divided into three main data subsets as follows: Split RatioThe dataset is partitioned into training, validation, and test sets using a standard ratio of 70:20:10. Experimental ConsistencyThis data partitioning strategy is applied consistently across all backbone model experiments to ensure fair and objective performance comparison between models. Cross-ValidationTo ensure model performance stability and to demonstrate that the results are not biased toward a specific data split, a 3-fold cross-validation scheme is also employed during model evaluation. Ethical Access Note Only the de-identified text benchmark is publicly released. Raw counselling audio is not shared due to privacy and ethical restrictions. All counselling data were collected with informed consent and ethical approval, and personally identifiable information has been removed or masked. This resource is not intended for direct clinical diagnosis or individual profiling.

Licenses

Similaires