Artificial Intelligence Applications in Thermal Power Plant Boiler Optimisation: Bibliometric and Topic Modelling Dataset (2014–2025)
Description
This dataset contains 322 Scopus-indexed publications on artificial intelligence applications in thermal power plant boilers and related thermal and energy-system applications, covering English-language journal articles and reviews published between 2014 and 2025. It supports the bibliometric and topic-modelling analysis reported in the associated study. The dataset includes raw bibliographic records, processed text, a document–term matrix, LDA outputs, topic-model evaluation results, temporal topic distributions, topic-stability results, and supporting bibliometric data. The selected seven-topic model includes document–topic probabilities, topic–term probabilities, topic labels, top terms, document topic assignments, and annual topic probabilities. The data can be used to reproduce the reported analysis, examine publication and collaboration patterns, investigate thematic development over time, and conduct further bibliometric or topic-modelling studies. Files are provided mainly in CSV format, with fitted R models supplied as RDS files
Files
Steps to reproduce
The analysis used the final 322 eligible Scopus publications comprising English-language journal articles and review papers published between 2014 and 2025. The title, abstract, and author-keyword fields were combined to form the text corpus. Text preprocessing included conversion to lowercase, removal of punctuation, numbers, and stopwords, treatment of domain-specific terms, and filtering of terms occurring fewer than five times, and 64 domain-specific stopwords were used. The resulting document–term matrix contains 322 documents and 1,676 terms. LDA was applied using Gibbs sampling. Candidate models with 2–10 topics were evaluated using 5,000 iterations, a burn-in of 2,000 iterations, thinning of 100, α = 50/k, β = 0.1, and seed 1234. Perplexity and probabilistic topic coherence were calculated for each candidate model. The seven-topic model was selected based on the evaluation results, topic interpretability, and intertopic structure. The selected seven-topic model provides document–topic probabilities in Main_322/LDA/theta_matrix.csv and topic–term probabilities in Main_322/LDA/beta_matrix.csv. Topic labels, top terms, topic summaries, and related outputs are provided in the Main_322/LDA/ directory. The fitted LDA models are also provided as RDS files. Topic stability was assessed by re-estimating the seven-topic model using five random seeds: 1234, 2024, 4321, 5678, and 9999. Topic matching and pairwise similarity results are provided in topic_stability_pairwise.csv and topic_stability_summary.csv. The corresponding fitted models are provided in lda_seed_models.rds. Temporal analysis was performed using document-level topic probabilities aggregated by publication year. The resulting annual topic distributions and topic-evolution data are provided in annual_topic_distribution.csv and topic_evolution_data.csv. The dataset includes supporting bibliometric and validation files, together with Documentation/parameters.csv, which records the main analysis settings. Documentation/verify.R provides structural checks for the deposited dataset, while Documentation/checksums.csv provides SHA-256 checksums for file verification. The dataset can be used to reproduce and inspect the reported bibliometric and topic-modelling analysis and to conduct further analyses of research trends in AI applications for thermal power plants, boiler systems, and related energy systems.
Institutions
- Central University of TechnologyFree State, Bloemfontein