Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Scaling Arabic Medical Chatbots Using Synthetic Data: Enhancing Generative AI with Synthetic Patient Records

Domaine:

natural language processinghealthcare

Type de record:

paperdatasetmodel
Créateur:
AllAhmHamSha
Hôte:avatar
The development of medical chatbots in Arabic is significantly constrained by the scarcity of large-scale, high-quality annotated datasets. While prior efforts compiled a dataset of 20,000 Arabic patient-doctor interactions from social media to fine-tune large language models (LLMs), model scalability and generalization remained limited. In this study, we propose a scalable synthetic data augmentation strategy to expand the training corpus to 100,000 records. Using advanced generative AI systems ChatGPT-4o and Gemini 2.5 Pro we generated 80,000 contextually relevant and medically coherent synthetic question-answer pairs grounded in the structure of the original dataset. These synthetic samples were semantically filtered, manually validated, and integrated into the training pipeline. We fine-tuned five LLMs, including Mistral-7B and AraGPT2, and evaluated their performance using BERTScore metrics and expert-driven qualitative assessments. To further analyze the effectiveness of synthetic sources, we conducted an ablation study comparing ChatGPT-4o and Gemini-generated data independently. The results showed that ChatGPT-4o data consistently led to higher F1-scores and fewer hallucinations across all models. Overall, our findings demonstrate the viability of synthetic augmentation as a practical solution for enhancing domain-specific language models in-low resource medical NLP, paving the way for more inclusive, scalable, and accurate Arabic healthcare chatbot systems. Accepted in AICCSA 2025

Visit

arxiv.org

Tags

Computation and Language

Similaires

Uganda HIE Synthetic FHIR R4 Dataset: 10,000 Patient Records Conforming to Uganda HMIS Data StandardsUZIMA-DS AI-Ready Synthetic DataScaling Effectiveness of Zero-Shot Cross-Lingual Retrieval with Synthetic Code-Switched DataSynthetic Climate Data Generation Using Generative Adversarial Networks for Drought Modelling in Southern AfricaEnhancing Fingerprint Gender Classification Using VGG19 Transfer Learning with Image-Based Synthetic OversamplingUZIMA-DS/UZIMA-DS-AI-Ready-Synthetic-Data

Uganda HIE Synthetic FHIR R4 Dataset: 10,000 Patient Records Conforming to Uganda HMIS Data Standards

Uganda HIE Synthetic Dataset: 10,000 FHIR R4 Patient Records This dataset comprises 10,000 syntheti

UZIMA-DS AI-Ready Synthetic Data

The AI-Ready Synthetic data study is being conducted by the UtiliZing health Information for Meaning

Scaling Effectiveness of Zero-Shot Cross-Lingual Retrieval with Synthetic Code-Switched Data

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Synthetic Climate Data Generation Using Generative Adversarial Networks for Drought Modelling in Southern Africa

Enhancing Fingerprint Gender Classification Using VGG19 Transfer Learning with Image-Based Synthetic Oversampling

Acute class imbalance and feature degradation in biometric fingerprint datasets significantly degrad

UZIMA-DS/UZIMA-DS-AI-Ready-Synthetic-Data

The AI-Ready Synthetic Data study, part of the UZIMA-DS hub (AKU & University of Michigan, funded by