A Gold-Standard Dataset for Benchmarking Balinese Extractive and Abstractive Text Summarization
Description
1. Research Hypothesis and Data Scope The central hypothesis guiding the creation of the BaliSummarization dataset is that the development of robust Balinese text summarization models—both extractive and abstractive—requires a high-quality, large-scale, and genre-diverse gold-standard resource, which is currently non-existent for this low-resource language. This dataset addresses this gap by providing human-generated reference summaries, enabling direct benchmarking for sequence-to-sequence (abstractive) and sequence-labeling (extractive) summarization tasks. The data is categorized into six distinct genres: articles, speeches (as known as Pidarta), and various forms of Balinese narrative (as known as Satua Bali), ensuring comprehensive coverage of linguistic variations across textual domains. 2. Data Overview and Interpretation The dataset includes three distinct categories: article or news, formal speech, and traditional folklore. These categories provide document-summaries, with the first category comprising articles, the second category being formal speech, and the last category being folklore. he data demonstrates a significant variance in document length, ranging from an average of 12.6 sentences/208.7 words (articlebats) to 73.8 sentences/1108.9 words (satuabaliweb), necessitating summarization models capable of handling extreme input lengths. The quality of human-generated reference summaries was validated using metrics like ROUGE-1, ROUGE-2, ROUGE-L, BLEU, and Cosine Similarity (FastText-based Embedding). We incorporated two metrics to validate our extractive summaries, i.e Fleiss's Kappa and Krippendorff's Alpha. Both of them resulting almost perfect agreement in all categories. 3. Notable Findings for Abstractive IAA score All categories achieved extremely high Cosine Similarity scores (above 0.91 across all annotation phases), significantly exceeding the 0.5 threshold. This indicates that annotators maintained a near-perfect consensus on the core meaning, topic, and semantic content of the summaries, despite using different phraseology. The primary challenge observed was low lexical overlap in bigrams (ROUGE-2), particularly for the articlebats, articlemlft, and pidarta categories, which failed to meet the 0.2 threshold in the independent phase. This suggests high linguistic variability in sentence construction when abstracting short or formal texts. The long narrative texts (folklore) and the articlesuara successfully passed all five IAA thresholds in the independent annotation, confirming their high reliability and serving as the most robust reference data within the collection.
Files
Steps to reproduce
The BaliSummarization dataset was constructed through the following three primary stages: Data Acquisition, Data Normalization, Data Annotation, and IAA Validation. 1. Data Acquisition and NormalizationData Acquisition and Normalization Source texts were collected from two streams: Web-scraping from public Balinese language online archives (e.g., Balinese language websites and blogs). Digitalization of printed media, including traditional folklore and formal speeches. Digitalization Instrument: Printed materials were converted to digital text using Optical Character Recognition (OCR) techniques. The software tools utilized for this purpose include specialized software such as imagetotext, picturetotext, cardscanner, and online converters like online-convert. Normalization Protocol: All collected texts underwent a normalization process, including: (a) removal of punctuation artifacts introduced by OCR errors; (b) standardization of spelling inconsistencies; and (c) sentence tokenization using a rule-based Balinese language tokenizer. 2. Annotation and Gold-Standard Generation Annotators: Three professional annotators with native proficiency in Balinese language and strong linguistic backgrounds were recruited for the summarization task. Extractive Summaries: Generated by all three annotator selecting a subset of original sentences from the source document to form the summary, ensuring factual consistency. Abstractive Summaries (Reference): Generated by all three annotators creating entirely novel, non-extractive summaries for the same set of source documents. Annotation Phases: The abstractive annotation adhered to a three-phase workflow: Pre-annotation (initial training), Pilot Annotation (initial testing and protocol refinement), and Independent Annotation (final data generation). 3. Inter-Annotator Agreement (IAA) Validation Protocol: IAA was calculated by comparing the abstracts written by the three annotators pairwise across all categories. Metrics and Thresholds: The assessment used five distinct metrics, with the following pass thresholds established for validation: Lexical Overlap: ROUGE-1 (>0.3), ROUGE-2 (>0.2), ROUGE-L (>0.2), and BLEU score (>0.2). Semantic Similarity: Cosine Similarity based on FastText Embeddings (>0.5). Software and Workflow: The IAA computation was executed using Python programming language, leveraging standard NLP libraries for ROUGE and BLEU calculation, and the FastText embedding model for semantic similarity computation.
Institutions
- Universitas UdayanaBali, Bukit Jimbaran
- Institut Teknologi Sepuluh NopemberJawa Timur, Surabaya
Categories
Funders
- Udayana UniversityBali, IndonesiaGrant ID: Program Bantuan Operasional Perguruan Tinggi Negeri Program Penelitian with contract no B/335-39/UN14.4.A/PT.01.03/2025