Authorship Attribution in Bahasa Indonesia

Published: 4 February 2026| Version 1 | DOI: 10.17632/gn9jsgkmrj.1
Contributor:
Yohan Muliono

Description

The dataset consists of textual data curated for authorship attribution research in the Indonesian language. It includes two distinct text types: short-text and long-text, representing different writing contexts and stylistic characteristics. Short-text data were collected from Twitter and reflect informal, concise, and spontaneous writing styles, while long-text data were obtained from Indonesian online news articles and represent formal, structured, and editorial writing.

Files

Steps to reproduce

Short-text data were collected when the official Twitter API was still publicly available, prior to the platform rebranding to X. Author accounts were identified through community input and public discussions on social media, targeting vocal and active writers to ensure sufficient data per author for authorship attribution. Data were retrieved using the official Twitter API. Long-text data were collected by web crawling Indonesian news articles from the Kompas news portal using Python-based scripts. Word_Count and Character_Count included in the data because in several research in Authorship Attribution, they were used as one of the feature

Institutions

Categories

Computer Science, Natural Language Processing, Attribution

Licence