Signals of Stance: A Dual Corpus of Reporting Verbs by Indonesian Academics in STEM and Humanities

Published: 3 July 2026| Version 1 | DOI: 10.17632/9cgt3sbv69.1
Contributors:
Muhammad Fikri Hasani, Clara Herlina Karjo, Menik Winiharti, Wishnoebroto Wishnoebroto

Description

This dataset consists of sentence-level extractions of reporting verbs (RVs) and their contextual sentences from academic papers written by Indonesian scholars (non-native English speakers). The corpus is divided into two files representing different academic disciplines: a Humanities dataset containing 192 sentences from 16 unique source articles, and a Science dataset containing 237 sentences from 22 unique source articles. Both CSV files share a uniform structure featuring five data columns: file_name (the source PDF filename), category (the semantic classification of the reporting verb: Argue, Show, Find, or Think), original_word (the exact verb form used in the text), base_verb (the lemmatized base form of the reporting verb), and sentence (the complete containing sentence). The Humanities corpus features a total of 192 extracted sentences with sentence lengths ranging from 11 to 70 words, averaging 30.95 words per sentence (median of 31 words). Semantically, the reporting verbs are distributed as follows: 117 sentences categorized as Argue (60.9%), 49 as Show (25.5%), 17 as Find (8.9%), and 9 as Think (4.7%). It contains 30 unique base verb lemmas. The top ten most frequent lemmas in this corpus are suggest (22 instances), write (15), reveal (15), report (15), indicate (13), show (13), state (11), argue (9), explain (9), and predict (8). The Science corpus features a total of 237 extracted sentences with sentence lengths ranging from 8 to 102 words, averaging 32.68 words per sentence (median of 29 words). Semantically, the reporting verbs are distributed as follows: 105 sentences categorized as Argue (44.3%), 95 as Show (40.1%), 20 as Find (8.4%), and 17 as Think (7.2%). It contains 23 unique base verb lemmas. The top ten most frequent lemmas in this corpus are show (38 instances), report (37), demonstrate (24), suggest (20), indicate (18), know (13), establish (13), maintain (12), reveal (11), and state (10). Lexically, there are 22 reporting verb lemmas shared between both corpora (including accept, acknowledge, conclude, demonstrate, establish, explain, hold, indicate, know, maintain, report, reveal, show, state, and suggest), with a Jaccard similarity index of 0.710. The Humanities corpus contains 8 unique lemmas not found in the Science corpus (add, think, say, find out, claim, point out, argue, realize). The Science corpus contains 1 unique lemma not found in the Humanities corpus (discover).

Files

Institutions

Categories

Social Sciences, Linguistics, Natural Language Processing

Licence