Bengali Descriptive Symptoms-to-Disease Identification Dataset
Description
This dataset contains Bengali-language descriptions of medical symptoms paired with corresponding disease labels. The symptom descriptions are written in natural conversational Bengali, simulating how a patient might explain their condition to a doctor. They were constructed using publicly available sources such as medical reports, blogs, dictionaries, and forums. Disease labels were assigned by matching extracted symptom keywords to known conditions. The dataset is intended for use in Natural Language Processing (NLP) research, particularly in tasks such as: Symptom-to-disease classification, Medical dialogue systems, and Low-resource language healthcare NLP etc. Format: .xlsx file with two columns: 'Descriptive Symptoms in Bengali', and 'Disease Label'. Keywords: Bengali NLP, medical dataset, symptom description, disease classification, healthcare AI
Files
Steps to reproduce
a. Collect public medical sources (blogs, dictionaries, forums, reports). b. Extract symptom keywords relevant to diseases. c. Restructure into Bengali sentences in natural patient-style speech. d. Assign disease labels using keyword-based mapping. e. Organize the dataset into .xlsx format with two columns. N.B. i. No personal or patient-identifiable information is included. ii. Disease labels are derived from keyword-based mapping, mostly synthetic. iii. The dataset is intended for NLP research only, not for clinical diagnosis or medical decision-making.
Institutions
- Rajshahi University of Engineering and TechnologyRajshahi Division, Rajshahi