Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

kechemale/symptom-extraction-in-Amharic-and-Tigrigna-languages

Domaine:

natural language processinghealthcare

Type de record:

model
Créateur:
kec
Hôte:
# Multilingual Symptom Extraction (English + Amharic) ## Model Description This model is a fine-tuned **XLM-R-base** model for extracting symptoms from patient-generated texts in **English and Amharic**. It was developed as part of an MSc thesis at the Technical University of Munich, aiming to support AI-powered symptom extraction platforms in multilingual healthcare settings. The model performs **Named Entity Recognition (NER)** to identify symptom mentions in unstructured texts using the **BIO tagging scheme**. ## Intended Uses & Limitations **Intended uses:** - Automatic symptom extraction from patient-generated texts in English and Amharic. - Research in multilingual biomedical NLP. - Integration into AI diagnostic platforms for low-resource languages. **Limitations:** - Trained only on the datasets described in the thesis; performance on other domains may vary. - Not validated for real-world clinical deployment without further testing and ethical approval. ## Training Data - Unlabeled **English, Amharic, and Tigrinya** corpora collected from diverse sources for domain adaptation. - A combination of publicly available datasets and synthetically generated data, which were then labeled for fine-tuning. - Preprocessed with tokenization, normalization, and **Doccano BIO tagging**. ## Evaluation - **Metrics:** Precision, Recall, F1-score for symptom extraction. - Results and detailed evaluation are described in the MSc thesis: > Negash Desalegn. (2025). *Bridging the Linguistic Gap in Healthcare: A Multilingual AI Approach for Symptom Extraction in Low-Resource Languages*. Technical University of Munich.

Visit

github.com

Tasks

information extractionnamed entity recognition

Languages

AmharicTigrigna