Synthetic Arabic Data for Scene Text Recognition
Published: 3 September 2021| Version 2 | DOI: 10.17632/gfc32vndz8.2
The dataset consists of 50,000 cropped images with embedded Arabic text. The labels were generated from an Arabic words corpus, which consists of 15 thousand words. The second dataset was collected from Twitter Arabic hashtags and contins 100 cropped images. The datasets were used in our publuished paper "Detecting Spam Images with Embedded Arabic Text in Twitter".
Steps to reproduce
Please refer to our published paper and this repo (https://github.com/Belval/TextRecognitionDataGenerator) for the steps.
University of York
Optical Character Recognition, Arabic Language, Recognition, Twitter