BanglaDigit

Published: 11 June 2026| Version 2 | DOI: 10.17632/s2nr768mhr.2
Contributors:
,
,
,
,
,

Description

This is a high quality research benchmark dataset for handwritten Bengali (Bangla) digit recognition. It consists of two structurally consistent variants of 20,000 and 60,000 samples respectively (BanglaDigit-20k and BanglaDigit-60k). To prevent compression artifacts, the data was obtained from 10 native Bengali speakers through an independent session on a tablet interface interface with a capacitive stylus, then saved as an uncompressed PDF. Unlike the data sets that use a geometric bounding-box centering method, BanglaDigit adopts Intensity Center of Mass (CoM) alignment technique, which is a mathematically sound curation methodology aligning each numeral in accordance with the weighted pixel intensity distribution to remove the translation bias systematically embedded in the Bengali script. All of the images are gray scale png and normalized to 28×28 pixels. All the digits (০–৯) of Bengali language have been taken with 2000 base samples for each class in the 20k version of the dataset, and 6000 samples per each class in the 60k version, using a stochastic augmentation engine. This is a dataset designed to provide a repeatable and stringent standard for both optical character recognition (OCR) research in the South Asian region and for quick architectural prototyping and testing, and deployment of edge computing. The six CNN architectures tested, ShuffleNetV2, MobileNetV2, ResNet18, DenseNet121, EfficientNetB0 and WideResNet50 achieve state of the art accuracy of 99.95% and 99.86% on the 20k and 60k splits respectively with a fixed training budget of 10 epochs.

Files

Institutions

Categories

Computer Science, Image Processing, Machine Learning

Licence