Tamil Brahmi Stone Inscription

Published: 24 August 2026| Version 1 | DOI: 10.17632/2spz2js9jp.1
Contributor:
Sukumar M

Description

The “Tamil Brahmi Stone Inscription Dataset (TBSI)** is the first open-source benchmark dataset dedicated to the processing and analysis of degraded Tamil stone inscriptions. Due to the lack of publicly available benchmark datasets for this specific domain, this repository bridges a critical gap in digital heritage and computational epigraphy. The dataset is compiled from the official archives of the **Tamil Nadu Department of Archaeology (TNDA)** and other publicly available epigraphic resources. Ancient Tamil stone inscriptions are invaluable historical records, but their automated analysis is severely hindered by natural degradation over centuries. This dataset is specifically curated to evaluate and train robust Computer Vision (CV) algorithms such as image enhancement, text segmentation, and Optical Character Recognition (OCR)—under real-world, highly degraded conditions.

Files

Steps to reproduce

As there is no open-source benchmark dataset dedicated to degraded Tamil stone inscription image processing, images were collected from the official archives of the Tamil Nadu Department of Archaeology and other publicly available epigraphic resources. The Tamil Nadu Department of Archaeology documents 93 Tamil Brahmi inscriptions discovered across more than 30 archaeological sites in Tamil Nadu, forming an important source for this research. Images were acquired from various archaeological sites with different levels of degradation (erosion, biological growth, surface cracks, shadow, uneven illumination, and low-contrast inscriptions). The resulting dataset contains representative examples of the challenges faced in the analysis of ancient Tamil stone inscriptions and is thus suitable for evaluation of robust computer vision algorithms.

Categories

Archeology, Computer Vision, Image Processing, Epigraphy

Licence