📢 Publish Your Research for Free - Full APC Waiver, No Hidden Charges. Submit Your Article Today! Submit Now →
JImage

ARTICLE TYPE : RESEARCH ARTICLE

Published on :   08 Aug 2026, Volume - 2
Journal Title :   WebLog Journal of Biotechnology and Biomedical Engineering | WebLog J Biotechnol Biomed Eng | WJBTBE
Source URL:   weblog icon https://weblogoa.com/articles/wjbtbe.2026.h0807
Permanent Identifier (DOI) :   doi icon https://doi.org/10.5281/zenodo.21860321

Machine Learning-Based Classification of Tuberculosis Using Whole-Blood Gene Expression Profiles: A Multi Cohort Transcriptomic Study with SHAP-Based Interpretability

Adeel Asghar 1 *
1Islamia University of Bahawalpur, Pakistan

Abstract

Background: Tuberculosis diagnosis remains a practical challenge in high-burden settings, where available tests have well-known limitations around sensitivity, cost, and turnaround time. Whole blood gene expression profiling offers a non-sputum-based alternative approach, but most published machine learning models in this space rely on complex pipelines or proprietary datasets. This study asked a simpler question: using publicly available microarray data and standard classifiers, can we build something that works, generalizes, and makes biological sense.

Methods: We used GSE19491 as the primary dataset, which contained 417 whole-blood transcriptomic samples across three classes: TB, Other Infection, and Control. After median imputation, z-score normalization, and ANOVA-based feature selection retaining the top 500 probes, SMOTE was applied to the training set to address class imbalance. Four models were trained and compared: Random Forest, SVM, Logistic Regression, and XGBoost. Bootstrap resampling with 200 iterations generated 95% confidence intervals for accuracy and weighted F1-score. SHAP values were calculated for the best-performing model. Pathway enrichment was performed using Enrichr. External validation used GSE37250, an independent Malawian cohort of 537 samples.

Results: Logistic Regression achieved the best performance with an accuracy of 0.929 and a weighted F1-score of 0.926, with a 95% confidence interval for accuracy of 0.869 to 0.976 (Table 1). Random Forest and XGBoost tied at 0.881 accuracy. AUC-ROC for the Random Forest was 0.91 for Control, 1.00 for Other Infection, and 0.95 for TB (Figure 2). On external validation in GSE37250, overall accuracy was 0.583, with TB precision of 0.79 and Other Infection recall of 0.73 (Figure 7). The top SHAP predictors included FKSG30, LOC402221, PDCD10, COX15, ARHGEF1, ACO2, IDH1, and ADAR (Table 3, Figures 5 and 6). Pathway enrichment showed consistent enrichment of TCA cycle, oxidative phosphorylation, Rho GTPase signaling, and mRNA editing pathways across multiple gene set libraries (Table 4).

Conclusions: A Random Forest classifier trained on 500 ANOVA-selected genes from whole blood microarray data can distinguish TB from other infections with high accuracy in the primary cohort and with meaningful but reduced performance in an independent population. The genes driving the model connect to mitochondrial metabolism, macrophage signalling, and innate immune regulation, which are all pathways TB is known to interact with. These findings support the biological validity of the transcriptomic signature and suggest a targeted gene panel built around the top SHAP predictors could be worth developing and testing in a prospective clinical setting.

Keywords: Tuberculosis; Gene Expression; Machine Learning; Random Forest; SHAP, Transcriptomics; GSE19491; GSE37250; Whole Blood; Biomarker

Citation

Adeel Asghar. Machine Learning Based Classification of Tuberculosis Using Whole-Blood Gene Expression Profiles: A Multi-Cohort Transcriptomic Study with SHAP-Based Interpretability. WebLog J Biotechnol Biomed Eng. wjbtbe.2026.h0807. https://doi.org/10.5281/zenodo.21860321