This dataset contains Bengali-language descriptions of medical symptoms paired with corresponding disease labels. The symptom descriptions are written in natural conversational Bengali, simulating how a patient might explain their condition to a doctor. They were constructed using publicly available sources such as medical reports, blogs, dictionaries, and forums. Disease labels were assigned by matching extracted symptom keywords to known conditions. The dataset is intended for use in Natural Language Processing (NLP) research, particularly in tasks such as: Symptom-to-disease classification, Medical dialogue systems, and Low-resource language healthcare NLP etc.
Format: .xlsx file with two columns: 'Descriptive Symptoms in Bengali', and 'Disease Label'.
Keywords: Bengali NLP, medical dataset, symptom description, disease classification, healthcare AI