HistoriKorpus

Published: 4 June 2026| Version 1 | DOI: 10.17632/b6p3jgdx83.1
Contributors:
,
,

Description

HistoriKorpus is the first dataset of OCR output and its corresponding ground truth for historical Indonesian newspapers published during the post-independence era (1946–1965). The dataset was constructed from 101 pages of newspapers sourced from the National Library of Indonesia (Perpusnas) via the Khastara digital repository. The corpus consists of two primary components. First, 1,489 articles that were manually transcribed by Indonesian university students, including history majors, to ensure accurate interpretation of archaic orthographic conventions. Second, a parallel corpus of 5,290 excerpts, each pairing a raw OCR output with its verified manual transcription (ground truth), totaling approximately 212k tokens and 1,365k characters. The dataset captures two historical Indonesian orthographic systems — the Van Ophuijsen spelling (pre-1947) and the Soewandi/Ejaan Republik spelling (1947–1972) — which differ substantially from contemporary Indonesian orthography. OCR was performed using a dedicated pipeline consisting of DocLayout-YOLO for text block detection and Tesseract 4 for text extraction. Baseline evaluation yielded a Character Error Rate (CER) of 8.25% and a Word Error Rate (WER) of 34.22%. HistoriKorpus is intended to serve as a benchmark resource for training and evaluating OCR systems, post-OCR correction models, and broader NLP applications for historical Indonesian text, including named entity recognition, text summarization, sentiment analysis, and historical fact-checking.

Files

Institutions

Categories

History, Data Science

Funders

Licence