Modelos Preditivos na Vigilância da Tuberculose uma Abordagem Metodológica para Cura Versus Óbito como Desfecho
Description
This material presents a reproducible methodological pipeline for developing machine-learning models to predict tuberculosis treatment outcomes in São Paulo, Brazil. It uses secondary surveillance data from TBWeb/SP and SINAN-TB covering the period from 2012 to 2024. The original dataset comprised 89,182 records, which were processed to obtain 69,154 eligible cases classified as cure or death. Data preparation included variable selection, data cleaning, missing-value treatment, consistency checks, and exclusion of variables with high incompleteness. The analytical dataset was stratified into 70% training and 30% independent testing subsets, with SMOTE applied exclusively to the training data. Five predictive approaches were evaluated: Random Forest, XGBoost, XGBoost Pareto, CatBoost, and CatBoost Pareto. Model performance was assessed using accuracy, sensitivity, specificity, precision, F1-score, ROC curves, AUC, and confusion matrices. Model stability and uncertainty were examined through five-fold stratified cross-validation and 1,000 bootstrap iterations. Variable importance and Pareto analysis were used to identify the predictors contributing most substantially to model performance and improve interpretability. The framework provides a transparent and reproducible approach for tuberculosis mortality risk prediction and supports epidemiological surveillance and evidence-informed public health decision-making.
Files
Categories
Funders
- Coordenação de Aperfeicoamento de Pessoal de Nível SuperiorMinistry of EducationFederal District, Brasília