Authorship Attribution in Bahasa Indonesia
Description
The dataset consists of textual data curated for authorship attribution research in the Indonesian language. It includes two distinct text types: short-text and long-text, representing different writing contexts and stylistic characteristics. Short-text data were collected from Twitter and reflect informal, concise, and spontaneous writing styles, while long-text data were obtained from Indonesian online news articles and represent formal, structured, and editorial writing.
Files
Steps to reproduce
Short-text data were collected when the official Twitter API was still publicly available, prior to the platform rebranding to X. Author accounts were identified through community input and public discussions on social media, targeting vocal and active writers to ensure sufficient data per author for authorship attribution. Data were retrieved using the official Twitter API. Long-text data were collected by web crawling Indonesian news articles from the Kompas news portal using Python-based scripts. Word_Count and Character_Count included in the data because in several research in Authorship Attribution, they were used as one of the feature
Institutions
- Binus UniversityJakarta, Jakarta