
This dataset contains processed and derived data supporting the figures and analyses in the manuscript “Durable Tuberculosis Risk Prediction Using a Three-Gene Blood Signature.” All data were derived from publicly available transcriptomic datasets deposited in GEO and ArrayExpress under accession numbers GSE79362, GSE107994, GSE157657, GSE94438, E-MTAB-6845, GSE89403, GSE83456, GSE111368, GSE117769, and GSE72509. No raw sequencing data are included.
Specifically, GSE79362, GSE107994, GSE157657, GSE94438, and E-MTAB-6845 were used for tuberculosis progression and validation analyses; GSE89403 was used for PET/CT-associated tuberculosis treatment-response analyses; and GSE83456, GSE111368, GSE117769, and GSE72509 were used for disease-specificity analyses. PET/CT-derived SUVmax data were obtained from the publicly available supplementary dataset of Malherbe et al.
The files provide figure-level source data, sample-level metadata, gene expression values, calculated signature scores, PET/CT variables, and gene-level association statistics used to generate the reported figures, ROC analyses, and statistical results.
File descriptions:
1. Data table and metadata of PET_SUVmax in Fig. 1.xlsx
This file contains the sample-level PET/CT-associated transcriptomic data used in Figure 1 to assess relationships between blood transcriptional signatures and lesion metabolic activity. It includes GEO sample metadata columns such as title, geo_accession, status, submission_date, platform_id, sample characteristics, sequencing and processing metadata, and repository relation fields. It also includes TB treatment-response metadata including disease.state.ch1, mgit.ch1, sample_code.ch1, subject.ch1, tgrv.ch1, time.ch1, timetonegativity.ch1, tissue.ch1, treatmentresult.ch1, and xpert.ch1. The key analysis columns are participant_no, timepoint, age, and SUVmax, together with expression values or derived values for signature genes and scores.
2. Data table of PET_Progression_6m in Fig. 1.xlsx
This file contains gene-level statistics used in Figure 1 for the 0.7–6 month pre-diagnosis progression window. Columns include gene for Ensembl gene identifier, symbol for gene symbol, rho for the correlation coefficient with PET/CT SUVmax, p.value for the nominal correlation P value, and padj for the adjusted correlation P value. The columns p_progressor and padj_progressor provide nominal and adjusted P values for association with progression status within the 0.7–6 month window.
3. Data table of PET_Progression_12month in Fig. 1.xlsx
This file contains gene-level statistics used in Figure 1 for the 0.7–12 month pre-diagnosis progression window. The column structure is the same as the 6-month file and includes gene, symbol, rho, p.value, padj, p_progressor, and padj_progressor. These data support the comparison between genes associated with lesion SUVmax and genes associated with subsequent tuberculosis progression over the 12-month window.
4. Data table and metadata of Progression in Fig. 2 and 3.xlsx
This file contains the main sample-level progression dataset used for Figures 2 and 3. Metadata columns include Sample, Group, Age, Sex, Time to diagnosis, and Source, where Group indicates progressor or nonprogressor status, Time to diagnosis provides the interval to tuberculosis diagnosis, and Source identifies the cohort or dataset. The file also contains expression values or derived values for signature genes and calculated scores. This file supports ROC curve generation, AUC calculation, and performance comparisons across time-to-diagnosis windows and age strata.
5. Data table and metadata of Other Diseases in Fig. 4.xlsx
This file contains the sample-level disease-specificity dataset used for Figure 4. Metadata columns include Sample, Age, Sex, Disease, and Source. The remaining columns provide expression values or derived values for signature genes and calculated scores across tuberculosis and non-tuberculosis disease groups. These data support the evaluation of cross-disease specificity of the tested tuberculosis risk signatures.
Together, these files provide the processed sample-level and gene-level source data required to reproduce the key figures, ROC analyses, correlation analyses, and disease-specificity comparisons reported in the manuscript. Signature definitions and scoring formulas are provided in the manuscript and/or supplementary materials.