ACSA Subtitle Scoring Corpus: English–Indonesian Subtitling of Our Planet Migrations
Description
This dataset supports the manuscript "Inference Relevance and Audience Access in the Subtitling of Netflix Nature Documentaries from English to Indonesian" (Baharuddin et al., submitted to Discover Global Society, Springer Nature). It contains all primary data generated during an empirical study of English-to-Indonesian subtitle quality assessment using the Audience-Centred Subtitling Adequacy (ACSA) framework.
Files
Steps to reproduce
1. Source material Select subtitle units from the target audiovisual text. In this study, source material was Our Planet: Migrations (Season 2, Netflix, 2023). Four episodes were selected, one per scorer group. Subtitle units were identified directly from the Netflix Indonesian subtitle track during viewing. 2. Scorer training Divide scorers into independent groups of approximately ten participants. Train each group on the ACSA rubric (see ACSA_Rubric sheet) using sample subtitle units not included in the scored corpus. Training covers the three dimensions: Micro (processing and readability), Meso (pragmatic alignment), and Macro (accessibility and global coherence). 3. Scoring procedure Each group independently viewed their assigned episode with the Indonesian subtitle track active and scored selected units using the ACSA rubric. For each unit, members scored individually, then reached consensus through group discussion before recording the final score. Scores were entered on a 1–5 scale for each dimension, with a written rationale note per score. 4. Data recording Each group recorded: the English source subtitle, the Indonesian target subtitle, the timecode, and the three ACSA scores with rationale notes. The four group spreadsheets were consolidated into the Scored_Corpus sheet. Group_Summary and Score_Distribution statistics were calculated from the consolidated corpus using standard Excel functions. 5. Replication with a new corpus To apply ACSA to a different subtitle corpus, use the ACSA_Rubric sheet to train scorers, select subtitle units from the target text, and follow steps 3–4. For inter-rater reliability, a minimum 10–15% subsample should be scored independently by two groups, with agreement calculated using Cohen's kappa or percentage agreement.
Institutions
- University of MataramWest Nusa Tenggara, Mataram