A Dataset of Children's Handwritten Bengali Characters for Deep Learning-Based Character Recognition

Published: 1 September 2026| Version 1 | DOI: 10.17632/2fh285fb82.1
Contributors:
,
,
,

Description

This dataset contains 24,277 handwritten Bengali character images collected from school-going children for research on Bengali handwritten character recognition. The dataset was specifically developed to capture the natural variability of children's handwriting, including differences in character formation, stroke patterns, writing styles, distortions, and incomplete strokes. The dataset comprises 70 character classes, including Bengali basic characters, modifiers, and digits. The samples were collected using a structured grid-based writing format on A4-sized paper, with each writing box designed to accommodate a single character. The dataset is intended for developing and evaluating machine learning and deep learning models for Bengali handwritten character recognition, particularly in educational and child-oriented applications. Unlike existing publicly available Bengali handwriting datasets, this dataset focuses specifically on children's handwriting, providing a resource for investigating recognition challenges associated with early-stage handwriting development. The dataset is intended for research and educational purposes, including Bengali handwritten character recognition, image classification, transfer learning, deep learning, explainable artificial intelligence, and comparative evaluation of recognition models. It may also support research on children's handwriting variability and computer-assisted educational technologies. However, the dataset is collected from a specific population of school-going children and therefore may not fully represent Bengali handwriting across all age groups, geographical regions, educational backgrounds, or writing environments. Consequently, researchers should consider these factors when assessing the generalizability of models trained or evaluated using this dataset.

Files

Steps to reproduce

The dataset can be reproduced and used for character-recognition experiments by following these general steps: Collect handwritten Bengali character samples from the target participants using a standardized writing format. Digitize and organize the collected samples according to their corresponding character classes. Preprocess the images through resizing, normalization, and other required image-processing operations. Divide the dataset into training, validation, and testing subsets. Apply appropriate augmentation techniques to the training samples. Train the selected machine learning or deep learning model using the specified experimental settings. Evaluate the trained model on the independent test set using standard classification metrics. Analyze the predictions using confusion matrices and misclassified samples to identify recognition patterns and errors. For explainability analysis, apply an appropriate XAI technique such as Grad-CAM to visualize the image regions influencing model predictions. This procedure provides a general framework for reproducing the dataset preparation, model development, and evaluation process while allowing researchers to adapt the specific model and experimental configuration to their requirements.

Institutions

Categories

Educational Technology, Optical Character Recognition, Applied Computer Science, Handwriting, Deep Learning

Licence