Image-to-Text Bilingual Dataset from Medical Prescriptions

Published: 19 January 2026| Version 2 | DOI: 10.17632/tg2hm7n2bs.2
Contributors:
,
,
,
,
,
,

Description

The 1000 handwritten images of medical prescriptions in the dataset are annotated in both Bangla and English, making it suitable for OCR, NLP, machine translation, and machine learning-based healthcare research. `annotations.csv` is a CSV file containing individual rows for each prescription and their respective Bangla and English texts. Those prescriptions that were clear and full were selected and the name and other identifying features of the patients blacked out. The images were obtained from hospitals, clinics and pharmacies in Gazipur, Bangladesh. The intention is research and education. Please see README.pdf within the dataset for more information on using it.

Files

Steps to reproduce

• Collect handwritten prescriptions from medical sources. • Scan or image each prescription in high resolution with a scanner or camera. • Type the Bangla and English text appearing in the pictures. • Images and text files should be named appropriately. • Use .jpg or .png for images and .csv format for text annotations.

Institutions

  • Bangabandhu Sheikh Mujibur Rahman Digital University

Categories

Text Extraction, Image Database

Licence