Digital Corpus Linguistics Based on Natural Language Processing (NLP) for English Islamic History Terminology in Indonesia
Description
Abstract: Indonesia produces a significant and growing body of English-language academic writing on Islamic history. However, the terminology carrying this scholarship, words such as ulama, pesantren, tarekat, and sultanate, has rarely been examined through the kind of systematic, corpus-based methods that linguistics currently has at its disposal. To address the gap, this study combined AntConc, a long-established concordance tool, with Orange Data Mining, a machine-learning platform, within a convergent mixed-methods design built around a corpus of 100 English-language articles on Indonesian Islamic history. At this reported stage, the corpus had been fully assembled and characterized. The corpus proved markedly concentrated in Java and Sumatra, which together supplied 96 of the 100 articles, in publications from 2025-2026, which accounted for two-thirds of the corpus, and in a small number of specialist journals, with just two venues supplying 38 percent of all articles. Ten recurring thematic strands were identified at the title level, led by historiography and intellectual thought. The quantitative and interpretive strands of the study were developed in parallel and are to be brought together through a triangulation procedure comparing AntConc's frequency and collocation findings with Orange's clustering and topic-modelling outputs. This study offered a transparent methodological template for pairing concordance-based and computational approaches to specialized religious-historical vocabulary, of likely interest to corpus linguists, Islamic studies scholars, and digital humanities researchers working on terminology. Keywords: Digital Corpus Linguistics, Natural Language Processing, AntConc, Orange Data Mining, Islamic History Terminology