Development of an Adaptive Text Summarization System Based on the T5 Transformer Model for Students with Dyslexia

Published: 11 June 2026| Version 1 | DOI: 10.17632/zv9d4x5vsp.1
Contributors:
,

Description

This dataset contains 110 Indonesian text samples developed to support research in automatic sign language translation, text simplification, and natural language processing. The data consist of original Indonesian texts accompanied by three levels of simplified target texts, namely Easy, Medium, and Hard. The Easy version uses simple vocabulary and short sentence structures to facilitate understanding, while the Medium version maintains more contextual information with moderate simplification. The Hard version preserves the original meaning and sentence structure with minimal modifications. The dataset was manually compiled and annotated to simulate linguistic transformations commonly required in sign language translation systems, where complex sentence structures are simplified while maintaining semantic accuracy. Each record includes an identifier, original text, simplified target texts, source page information, and relevant keywords. This dataset is intended for applications such as text simplification, machine translation, sign language translation, accessibility technologies for deaf and hard-of-hearing communities, and the development of transformer-based language models including T5, mT5, BART, and IndoBART. The dataset is provided in Microsoft Excel (.xlsx) format and serves as a valuable resource for researchers working in Indonesian language processing, educational technology, and assistive communication systems.

Files

Steps to reproduce

The dataset was developed to support research on adaptive text summarization for students with dyslexia. First, Indonesian reading texts were collected from educational materials and publicly available learning resources. The collected texts were manually reviewed and cleaned to remove duplicates and irrelevant content. Each text was then simplified into three levels of complexity: Easy, Medium, and Hard. The Easy level contains highly simplified sentences with basic vocabulary and shorter structures. The Medium level preserves more contextual information while maintaining readability. The Hard level retains the original text with minimal modifications. Keywords and source page information were added to each record. The final dataset was stored in Microsoft Excel (.xlsx) format and prepared for training and evaluation of Transformer-based models such as T5.

Categories

Natural Language Processing, Transformer-Based Deep Learning, Transformer Language Model

Licence