SaSamBo Corpus Dataset

Published: 9 March 2026| Version 2 | DOI: 10.17632/f2jp385zkr.2
Contributors:
, Sri Wahyuni

Description

raw data from SaSamBo Corpus https://korpussasambo.id/

Files

Steps to reproduce

1. Collect Indonesian–Mbojo, Indonesian–Samawa, and Indonesian–Sasak speech and text data from publicly available sources and manually curated regional corpora. 2. Perform speech transcription using an automatic speech recognition (ASR) pipeline and extract textual content from scanned documents using OCR when applicable. 3. Clean and normalize all texts by removing punctuation, non-linguistic symbols, duplicated entries, and inconsistent spelling. All data are anonymized to remove personally identifiable information. 4. Align sentence pairs manually and semi-automatically to form parallel corpora for each language pair. 5. Encode sentences using multilingual Transformer models (mBERT, DistilBERT, XLM-R Base, XLM-R Large, and LaBSE). 6. Apply LSTM-based temporal smoothing to the embedding sequences to obtain final semantic representations. 7. Evaluate lexical alignment using Top-1 Accuracy, Top-K Accuracy, BLEU, Mean Reciprocal Rank (MRR), and BERTScore metrics. 8. Split the dataset into training, validation, and test sets using a 70:15:15 ratio. All experiments were conducted using Python and PyTorch on NVIDIA GPU hardware. The provided dataset contains the final aligned text pairs and evaluation-ready samples.

Categories

Text Processing, Text Mining, Text Segmentation

Licence