Thai–English Multiscript Text Image (TEMS) Dataset

Published: 13 July 2026| Version 5 | DOI: 10.17632/ntdmgksh9w.5
Contributors:
, olarik surinta

Description

The Thai-English Multiscript Text Image (TEMS) Dataset is introduced as a publicly available resource for multilingual scene text recognition and related computer vision research. The dataset comprises 5,000 cropped text images systematically extracted from 1,625 natural scene photographs captured using smartphone cameras. Text regions were collected from diverse real-world environments, including billboards, commercial storefronts, road signs, menus, packaging, and publication covers, representing realistic visual conditions commonly encountered in practical applications. The dataset contains Thai text, English text, mixed Thai-English text, numerals, punctuation marks, and special symbols. In total, the TEMS Dataset includes 161 unique character classes, comprising Thai script characters (44 consonants, 20 vowels, 4 tone marks, 3 punctuation marks, and 10 Thai numerals), English uppercase and lowercase letters (52 characters), Arabic numerals (10 digits), and 18 special symbols. The cropped text images exhibit substantial variation in text length, font style, illumination conditions, background complexity, and image dimensions, ranging from 45 to 1,512 pixels in width and from 18 to 275 pixels in height. These characteristics provide realistic and challenging conditions for multilingual scene text analysis and recognition tasks.

Files

Steps to reproduce

1. Original Image Acquisition and Character Analysis The dataset originated from 1,625 natural scene photographs captured using smartphone cameras in real-world environments. Analysis of the collected images identified 161 unique character classes, including Thai consonants, vowels, tone marks, punctuation marks, Thai numerals, English uppercase and lowercase letters, Arabic numerals, and special symbols. 2. Text Region Extraction and Standardization A total of 5,000 cropped text images were generated from the 1,625 source photographs through a manual text extraction process. Each cropped image corresponds to a text region containing Thai text, English text, mixed Thai-English text, numerals, punctuation marks, or special symbols. All images were stored in JPEG format. Annotation files generated during data preparation, such as XML files, are not included in the public dataset release. 3. Dataset Labeling through Filename-Based Annotation To maintain a simple and consistent dataset structure, transcription labels were embedded directly into the image filenames. Each filename follows a standardized naming convention consisting of a unique image identifier and the corresponding ground-truth text label. This approach enables users to associate each image with its transcription without requiring separate annotation files.

Institutions

Categories

Handwriting Recognition, Image Classification, Deep Learning

Funders

Licence