Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

StrokeID-NER : A Stroke Name Entity Recognition Dataset for Indonesian Patient-Generated Health Queries

Domaine:

natural language processinghealthcare

Type de record:

dataset
Créateur:
CorYasMozKim
Éditeur:
Zenodo
Hôte:avatar

Stroke is the leading cause of death and disability in Indonesia, yet clinical NLP tools for stroke triage support remain limited by the absence of domain-specific annotated resources for patient-generated health text. We introduce StrokeID-NER, a stroke-specific named entity recognition (NER) dataset for informal Indonesian patient queries, comprising 1,742 queries annotated with 6,765 entity spans across four clinically grounded entity types — Symptom, Diagnosis, Temporal, and Risk Factor — with inter-annotator agreement (κ = 0.857). We fine-tuned and evaluated five NER systems on our corpus, spanning monolingual and multilingual transformer models at two scales and a GPT-5.4 zero-shot baseline.  Fine-tuned XLM-RoBERTa-large achieved the best performance (macro F1 = 0.7823), outperforming the zero-shot baseline by 11.4 points, with the largest advantage concentrated in Symptom and Risk Factor entities requiring recognition of informal Indonesian, Javanese regional terminology, and clinical context. Hapax rates of 72.0–86.9% across entity types and 68.3% cross-model false negative overlap indicate that performance gains are constrained by lexical coverage rather than model architecture, suggesting vocabulary expansion as the most effective path to improvement. StrokeID-NER, the annotation guidelines, and all trained models are released publicly at https://doi.org/10.5281/zen….

Visit

doi.org

Tasks

information extractionnamed entity recognition

Languages

Ndasa

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

An Ontology-based Name Entity Recognition NER and NLP Systems in Arabic Storytellingnardoki/Amharic-Name-Entity-RecognitionNLP: Rule Based Name Entity RecognitionDataset for "A Benchmark Suite for Indonesian Mental Health NER: Combining Real Counselling Audio Transcripts and Adapted Text Data"NERAMazigh: A Named Entity Recognition Dataset for the Amazigh LanguageELNER-DZ: A Dataset for Named Entity Recognition and Entity Linking in Algerian Arabic Dialect

An Ontology-based Name Entity Recognition NER and NLP Systems in Arabic Storytelling

nardoki/Amharic-Name-Entity-Recognition

### 🇪🇹 Amharic Named Entity Recognition (NER) A transformer-based Amharic Named Entity Recognition

NLP: Rule Based Name Entity Recognition

Named Entity Recognition (NER) is an information extraction task aimed at identifying and classifyin

Dataset for "A Benchmark Suite for Indonesian Mental Health NER: Combining Real Counselling Audio Transcripts and Adapted Text Data"

Description This dataset focuses on utilizing primary data in the form of original Indonesian psych

NERAMazigh: A Named Entity Recognition Dataset for the Amazigh Language

NERAMazigh is a manually annotated Named Entity Recognition (NER) dataset for the Amazigh language,

ELNER-DZ: A Dataset for Named Entity Recognition and Entity Linking in Algerian Arabic Dialect

ELNER-DZ is the first large-scale dataset for Named Entity Recognition (NER) and Entity Linking (EL)