Narrative Tone, Evaluative Polarity, and Legitimacy in BBC Monitoring’s Coverage of the Belt and Road Initiative

Published: 21 April 2026| Version 1 | DOI: 10.17632/8zrtz8h7kv.1
Contributor:
Tariq H. Malik

Description

Dataset. This study draws on an original dataset of 1,869 BBC Monitoring reports related to China’s Belt and Road Initiative (BRI) from 2013 to 2024. BBC Monitoring is an appropriate source for the present analysis because it functions as a specialist media-monitoring and translation service that curates foreign media content for institutional consumers, including governments, researchers, journalists, and policy organisations, rather than for a mass public audience. Its reports are derived from foreign newspapers, websites, and social media across more than 100 languages and are reformulated through processes of summary, translation, and contextual interpretation. The dataset therefore captures not merely raw reporting on the BRI, but an institutionalised layer of mediated narrative reconstruction. Substantively, the corpus spans economic, geopolitical, environmental, and cultural dimensions of the BRI. Each record includes the document ID, date, title, abstract, region, and text length, which together provide the basis for subsequent coding and quantitative text analysis. To construct the coding scheme, 100 reports were first manually coded to identify tone, themes, and narrative framing, and this protocol was then scaled to the full corpus through automated coding procedures.

Files

Steps to reproduce

Steps to reproduce. First, compile a corpus of 1,869 BBC Monitoring reports on the Belt and Road Initiative (BRI) from 2013 to 2024. For each report, record the document ID, publication date, title, abstract, region, and word count in a spreadsheet. BBC Monitoring is appropriate because it translates, summarizes, and contextualizes foreign media for institutional users, allowing analysis of mediated narrative reconstruction rather than raw mass-public news. Second, develop the coding framework through a mixed manual and automated process. Begin with a manual coding subsample of 100 reports to identify tone, themes, and narrative framing. Use this pilot stage to define the core categories, especially neutral narrative and developmental narrative, and then extend the coding protocol to the full corpus using automated coding. Binary variables are coded as 1 when the feature is present and 0 otherwise. Third, preprocess the text for NLP analysis. Tokenize each report into words or phrases and apply a lexicon-based approach to generate document-level measures of subjectivity and polarity. Subjectivity is scored on a 0 to 1 scale, where higher values indicate more opinionated or emotional language. Polarity captures evaluative tone as positive, neutral, or negative, operationalized in the study as 1, 0, and -1. Fourth, estimate the models in stages. Start with descriptive statistics and correlation analysis, then run robust regression to test the association between subjectivity and polarity while reducing the influence of outliers. After that, estimate structural equation models (SEM) to assess the direct and indirect paths from neutral and developmental narratives to polarity through subjectivity, first without and then with rhetoric and framing controls. Finally, interpret the results in line with the article’s stated epistemology. The analysis is presented not as a search for universal causal laws, but as an examination of coherence within a single institutional narrative system. Thus, model fit and high explained variance are treated as evidence of narrative-system closure and editorial consistency within BBC Monitoring rather than as externally generalizable prediction.

Institutions

Categories

International Political Economy, China, Legitimacy Theory, Digital Media, Belt and Road Initiative

Licence