Bridging the gap between academic exploration and industrial reality in supercritical fluid extraction: a machine learning and kinetic approach

Published: 8 December 2025| Version 1 | DOI: 10.17632/bzfs5xz873.1
Contributors:
,
,
,

Description

Abstract of the related article: Supercritical Fluid Extraction (SFE) has evolved from a niche technique to a pivotal platform for the circular bioeconomy. This study presents a critical, data-driven assessment of the SFE landscape, integrating text mining, kinetic modeling, and machine learning (Random Forest) to analyze 1,204 scientific articles (2010–2024) and 3,844 patents. The results reveal a "science-push" phenomenon, a structural shift towards microalgae and waste valorization, and a strategic divergence between academic exploration and industrial intellectual property protection. Description of the data files: This dataset contains the processed data and scripts used to generate the results and figures presented in the article. 1) Artigos_SFE_FINAL_COMPLETO.xlsx: The harmonized bibliographic dataset of 1,204 scientific articles. It includes standardized columns for Matrix Categories (e.g., Marine Sources, Agro-Waste) and Target Compounds, cleaned via text mining algorithms. 2) Tabela_Cinetica_K.xlsx: The output of the kinetic modeling analysis, containing the calculated growth rate constants (k), Standard Errors, and R2 values for all studied matrix and target categories (2010–2024). 3) Scripts_R_Analysis.R: The complete R programming script used for: a) Data harmonization and cleaning. b) Kinetic modeling (non-linear regression). c) Multivariate analysis (Random Forest and Negative Binomial Regression). d) Generation of figures and patent landscape plots. Methodology: Data were retrieved from Scopus, Web of Science, and PubMed (scientific articles) and Lens.org (patents). Text mining and statistical analyses were performed using R (v. 4.4.1) with packages dplyr, stringr, minpack.lm, and randomForest.

Files

Steps to reproduce

To reproduce the analysis and figures presented in the manuscript: 1) Software: Ensure you have R (version 4.0 or higher) and RStudio 2) installed.Files: Download all files from this dataset (.xlsx, .csv, and .R) and save them into a single local directory. 3) Execution: Open the file Scripts_R_Analysis.R in RStudio. 4) Setup: The script contains lines to automatically install and load the necessary packages (dplyr, ggplot2, randomForest, etc.). 5) Run: Execute the code blocks sequentially. The script reads the harmonized Excel file (Artigos_SFE_FINAL_COMPLETO.xlsx) and generates the statistical models (Kinetic, Random Forest) and high-resolution figures used in the article.

Institutions

  • Universidade do Estado de Santa Catarina

Categories

Chemistry, Chemical Engineering, Food Science, Data Science

Funders

Licence