Source-linked AI summary
Population Health-Based Machine Learning Reveals Associations Between Psychosocial Factors and Chronic Kidney Disease
Md. Atik Shams, David Eisenberg, Sumaiya Fatema, Asma Sultana, D. M Hasibul Islam, Junnatul Mawa, Anindita Datta, Nafiya Ahmed, Danastan Tasaouf Mridula, SK. Sazid Mahmud, Simon Bin Akter, Tanjila Helaly, Jorge Fresneda Fernandez, Humayera Islam, Tanmoy Sarkar Pias
TL;DR
CKD can remain undiagnosed, while population surveys offer broader coverage but introduce missingness, heterogeneity, and class-imbalance challenges. This study combines advanced preprocessing with a customized stacked ensemble to classify self-reported CKD and uses SHAP to characterize associated features. The framework achieved statistically significant improvements over single Gradient Boosting and Logistic Regression models and was presented as an interpretable, cost-efficient tool for large-scale CKD classification.
Problem
Hospital-based databases primarily capture patients already receiving treatment, leaving broader population risk and early-stage CKD insufficiently represented in available data.
Method
The study uses BRFSS and NHIS survey data, multiple imputation and sampling methods, a customized stacked ensemble, and SHAP-based interpretation for CKD classification.
Results
The proposed ensemble showed statistically significant improvements over single Gradient Boosting and Logistic Regression on both NHIS and BRFSS cohorts.
Takeaways & Limitations
The framework provides an interpretable, cost-efficient approach for large-scale CKD classification and helps identify features associated with kidney-disease risk.
Takeaways & Limitations
The study relies on self-reported CKD rather than laboratory-confirmed CKD.
Abstract
from arXiv · showhide
Chronic kidney disease (CKD) progresses silently and severely undermines quality of life, making early detection critical for improving patient outcomes. We present a two-part study that combines large-scale telehealth data with advanced machine learning to both classify self-reported CKD status and identify key drivers of disease. Using selected features from the Behavioral Risk Factor Surveillance System (BRFSS 2021: 438,693 samples; BRFSS 2019: 418,268 samples) and the National Health Interview Survey (NHIS 2021: 29,482 samples; NHIS 2020: 31,568 samples), we addressed missing data with nine state-of-the-art imputation methods and mitigated class imbalance via sampling strategies. Our customized stacked ensemble model achieved balanced accuracy of 72.56-76.12%, with corresponding AUROC scores of 79.59-82.29%. SHapley Additive exPlanations (SHAP) analysis, followed by clinical review, highlighted critical predictors, including regular medical check-ups, age, blood pressure, and indicators of mental health stress. These findings deliver a robust and interpretable framework for CKD risk stratification and provide actionable insights into its associated factors.
Authors
The paper lists its authors and their institutional affiliation in Bangladesh.
- The author list includes Md. Atik Shams and collaborators.The listed authors span multiple numbered affiliations.
- The authors include David Eisenberg, Sumaiya Fatema, and Asma Sultana.
- The author list identifies Tanmoy Sarkar Pias as a corresponding author.An asterisk follows the name.
2 Department of Computer Science, BRAC University, Dhaka 1212, Bangladesh
The listed Department of Computer Science affiliation is at BRAC University in Dhaka, Bangladesh.
- The affiliation is a Department of Computer Science.
- The institution is BRAC University.
- The affiliation is located in Dhaka 1212, Bangladesh.
4 Institute of Biological Sciences, Rajshahi University, Rajshahi 6205, Bangladesh
The listed affiliations include institutions in Bangladesh and the United States.
- The Institute of Biological Sciences affiliation is at Rajshahi University in Rajshahi, Bangladesh.
- Another affiliation is the Martin Tuchman School of Management at the New Jersey Institute of Technology in Newark, New Jersey, USA.
Introduction
CKD prediction in population surveys is motivated by the need to identify risk beyond hospital-treated populations while addressing missing data, heterogeneity, and class imbalance. The study therefore combines advanced imputation, resampling, ensemble classification, and interpretability methods for large-scale CKD risk stratification.
- Motivation: CKD is a major public-health concern, and hospital-based databases can miss early-stage cases because they primarily include patients already receiving treatment.
- Motivation: Cross-sectional BRFSS and NHIS surveys cover broader populations but contain noisy self-reported data and substantial missingness.
- Challenges: Missing values and class imbalance can reduce the performance of machine-learning models applied to health surveys.
- Research gap: Prior CKD studies often did not address mixed data types, missing data, and severe class imbalance simultaneously.
- Approach: The study evaluates multiple imputation methods and resampling strategies rather than relying on a single preprocessing approach.The introduction describes comparisons among advanced imputation and sampling techniques.
- Approach: A stacking ensemble combines XGBoost, Random Forest, Logistic Regression, and a multilayer perceptron for heterogeneous survey data.
Performance on NHIS datasets
Across the NHIS datasets, stacked ensembles and selected imputation–sampling combinations achieved the strongest reported classification results, with performance varying by metric and dataset.
- The Stacked Ensemble model achieved the highest balanced accuracy and AUROC on NHIS 2020.Table 1 ranked models by balanced accuracy.
- 0.7537 balanced accuracy, 0.7397 sensitivity, and 0.7700 specificity were achieved by ANN_2_Layer with GAN imputation and RUS sampling on NHIS 2021.Its cross-validation balanced accuracy was 0.6730 ± 0.01.
- 0.8069 AUROC was achieved by logistic regression with XGBoost imputation and ROS sampling on NHIS 2021.The cross-validation AUROC was 0.7776 ± 0.02.
- 0.8229 AUROC and 0.7661 specificity were achieved by the Stacked Ensemble with GAN imputation and RUS sampling on NHIS 2020.Its cross-validation AUROC was 0.7724 ± 0.02.
- 0.7648 balanced accuracy was achieved by GradientBoosting with GAN imputation and ROS sampling on NHIS 2020.ANN_2_Layer with XGBoost imputation and SMOTE sampling achieved the highest sensitivity of 0.7498.
Performance on BRFSS datasets
On BRFSS datasets, stacked ensembles repeatedly ranked among the strongest models, while SHAP analysis identified physical, demographic, clinical, and psychosocial factors associated with CKD classification.
- BRFSS 2021 performance: 0.7547 sensitivity was achieved by the BRFSS 2021 Stacked Ensemble with XGBoost imputation and RUS sampling.The same pipeline achieved AUROC 0.7960 and balanced accuracy 0.7256.
- BRFSS 2021 performance: 0.7977 AUROC and balanced accuracy were achieved by GradientBoosting with GAN imputation and RUS sampling on BRFSS 2021.The passage reports these as the highest AUROC and balanced accuracy for that dataset.
- General-population screening: The Stacked Ensemble model performed among the best models for early CKD screening across unadjusted general-population cohorts.Leading pipelines had PPVs of 8.5%–9.8% and high NPVs, with baseline CKD prevalence around 3.9%.
- Cross-dataset comparison: The Stacked Ensemble appeared among the five best models across all four test sets.GAN imputation with RUS sampling showed balanced performance in three datasets.
- Cross-dataset comparison: The proposed framework outperformed comparison models across the evaluated datasets.McNemar’s tests found significant improvements on NHIS 2020, including p<0.001 against Gradient Boosting with GAN+ROS.
- Clinical baseline comparison: Simple baseline models had sensitivity 0.000, whereas the proposed framework achieved sensitivity above 0.750 across datasets.The baselines were trained to predict the major Healthy class under prevalence ranging from 3.9% to 7.3%.
- Feature importance analysis using SHAP values: Difficulty walking, blood pressure, and age were the top BRFSS 2021 SHAP features, with values of 0.19, 0.15, and 0.14.Employment status and physical health also had notable values of 0.13 and 0.10.
- Feature importance analysis using SHAP values: Mental health, alcohol consumption, BMI, and parental separation also contributed to BRFSS 2021 classifications.Their reported SHAP values were 0.03, 0.05, 0.02, and 0.02, respectively.
Discussion
The study combines imputation, sampling, stacked ensembles, and SHAP interpretation to classify self-reported CKD and identify influential associated factors in large health surveys. The framework supports risk stratification, but its findings are bounded by self-reported outcomes, cross-sectional associations, unweighted cohorts, and the need for external validation.
- Interpretability: SHAP identified blood pressure, age, difficulty walking, employment status, income, physical health, and mental-health-related indicators as influential CKD-classification factors.Across the reported analyses, blood pressure and age repeatedly ranked among the strongest predictors, while other psychosocial and functional variables also contributed.
- Performance: Statistically significant improvements over single Gradient Boosting (p < 0.001) and Logistic Regression (p ≤ 0.015) were observed on both NHIS and BRFSS cohorts.McNemar's tests supported non-random improvements from combining multiple base estimators with a neural-network meta-learner.
- Method: The framework addresses missing values and class imbalance while combining eleven models, including a customized stacked hybrid architecture with an MLP meta-learner.The architecture combines XGBoost, Random Forest, and Logistic Regression through an MLP meta-learner to classify CKD.
- Limitations: The model classifies likelihood of self-reported diagnosis rather than true physiological onset because CKD labels were not laboratory-confirmed.The study relies on self-reported CKD instead of clinical records, eGFR, or albuminuria, and early-stage CKD is largely asymptomatic.
- Limitations: Cross-sectional associations may reflect comorbidities, frailty, healthcare utilization, or downstream illness consequences rather than upstream risk factors.The authors specifically identify difficulty walking, physical health, regular checkups, and difficulty concentrating as potentially proxy variables.
- Limitations: The unweighted survey analysis is not nationally representative, and broader generalizability requires validation in external cohorts.The authors also note that future work could apply survey weights and validate the configurations externally.
- Implications: The framework may help prioritize high-risk individuals for targeted CKD screening and confirmatory laboratory testing.The authors describe the approach as a low-cost population-survey-based triage system rather than a replacement for physiological diagnosis.
- Deployment: Inference latency of 0.018 to 0.025 milliseconds per patient supports compatibility with typical EHR systems without dedicated GPU infrastructure.The reported throughput exceeds 38,000 to 54,000 evaluations per second.
Methods
The study uses four national health-survey datasets, extensive feature preprocessing, multiple imputation methods, and resampling to classify self-reported CKD. The workflow addresses missingness, imbalance, heterogeneous survey variables, and leakage prevention.
- Missing-data handling: Nine imputation techniques, including GAN, diffusion, XGBoost, and autoencoders, were used to handle missing values in complex health-survey data.Categorical and ordinal values generated by continuous-output models were rounded to the nearest valid category.
- Datasets: Four BRFSS and NHIS survey datasets provide large-scale information on health behaviors, conditions, and socioeconomic factors.BRFSS 2021 and 2019 include 438,693 and 418,268 participants; NHIS 2021 and 2020 include 29,482 and 31,568 participants.
- Feature selection: Thirty-two input features cover health status, adverse childhood experiences, demographics, disability, behaviors, and socioeconomic characteristics.The BRFSS feature set contains 28 categorical and 4 numeric variables.
- Preprocessing: An 80/20 stratified train-test split was performed before imputation, with imputation algorithms fitted only on the training data.The target variable was excluded while missing values were inferred from independent features.
- Outcome definition: The primary outcome was self-reported CKD based on survey questions asking whether respondents had been told they had kidney disease or weak or failing kidneys.Kidney stones, bladder infections, and incontinence were excluded from the outcome definitions.
- Class balancing: Resampling methods addressed the 3.9% positive-case imbalance, producing approximately 50:50 class distributions for most methods.AdaSyn slightly improved minority-class representation, while ROS, RUS, and SMOTE generally produced near-balanced training data.
Predictive Modeling Architectures
The study evaluates eleven predictive approaches spanning classical machine learning, deep learning, and a heterogeneous stacked architecture. The proposed ensemble combines three base models with an MLP meta-learner, while SHAP provides prediction interpretation.
- Model evaluation: Eleven predictive modeling approaches were evaluated across BRFSS and NHIS datasets for systematic CKD classification.The models were trained on one survey year and evaluated on corresponding datasets, including cross-year evaluation.
- Model evaluation: The evaluated architectures include classical machine learning, deep learning, and a custom stacked hybrid model.The model selection targeted high dimensionality, class imbalance, and complex data patterns.
- Stacked ensemble: The stacked architecture combines XGBoost, Random Forest, and Logistic Regression base estimators with an MLP meta-learner.The meta-learner uses out-of-fold predictions from five-fold cross-validation, while the 20% test set remains untouched until final evaluation.
- Stacked ensemble: The proposed architecture differs from commonly used homogeneous ensembles and linear meta-learners by using heterogeneous base models and an MLP meta-learner.The paper identifies limited prior use of heterogeneous stacked ensembles for CKD classification.
- Interpretability: SHAP was used to quantify feature contributions and interpret the ensemble’s CKD classifications.The authors selected SHAP instead of heuristic tree metrics and approximation methods such as LIME.
Code availability
The study reports that its analyses used Python and Google Colab, with code publicly available on GitHub.
- Code availability: All analyses were conducted in Python on Google Colab, and the code is publicly available on GitHub.The repository URL is https://github.com/AtikShams/CKD.
Additional Information
The supplementary document accompanies the submission, and materials requests are directed to the corresponding author.
- Additional Information: The supplementary document is attached with the submission, with materials requests directed to Tanmoy Sarkar Pias.The passage identifies the correspondence contact for requests concerning materials.
Figures and tables
The figures and tables document the CKD classification workflow, compare model pipelines across survey datasets, and organize the clinical, demographic, activity, and ACE features used. They also describe the proposed stacked architecture and its evaluation tables.
- Figure 1 presents the complete workflow for CKD classification, from data collection through supporting healthier lives.
- Figure 2 compares five leading modeling pipelines for NHIS 2020 and NHIS 2021 using confusion matrices.The pipelines vary by classifier, imputation technique, and sampling strategy.
- Figure 3 evaluates calibration across NHIS 2020, NHIS 2021, BRFSS 2019, and BRFSS 2021 using GAN plus RUS pipelines.Curves closer to the diagonal indicate higher probabilistic calibration.
- Figure 4 compares AUROC, balanced accuracy, sensitivity, and specificity across four datasets, with Ensemble Stacking consistently among the top models.The figure combines AUROC curves, radar charts, and sensitivity–specificity scatter plots.
- Figure 5 ranks feature contributions, highlighting blood pressure, difficulty walking, age, and income level across BRFSS and NHIS datasets.Longer SHAP bars indicate stronger contributions to individual classifications.
- Figure 6 shows a heterogeneous stacked ensemble using XGBoost, Random Forest, and Logistic Regression as base estimators and MLP as the meta-learner.
NHIS 2020 Feature map
The NHIS 2020 feature map documents demographic, activity, disability, health, behavioral, socioeconomic, and psychosocial variables used in the analysis. Supporting materials describe their distributions, coding, and relationships discussed in CKD context.
- The feature map covers demographic, activity-related, disability, and health variables, including BMI, vision, walking difficulty, insurance, mental health, blood pressure, and cholesterol awareness.
- Ordinal survey variables encode graded responses such as difficulty from no difficulty to inability, while BMI categories range from underweight to obese.
- The distribution plot describes 32 features spanning demographics, adverse childhood experiences, and behavioral factors, separated into positive and negative kidney-disease cases.
- The study discusses brain-regulated neurotransmitters and hormones, including dopamine and serotonin, as pathways that may affect renal function.
- CKD is reported as more common in older adults, with prevalence of 34% among those over 65 versus 12% among those aged 45–64 years, although aging has no direct linear relation to CKD development.
- The discussion links psychological stress, parental separation, alcohol use, education, housing, and employment-related health effects with CKD-associated factors, while preserving indirect or qualified relationships.