Abstract
In this study, in silico methodologies were applied to publicly available gene expression datasets to identify potential biomarkers of prostate cancer. Current diagnostic methods, including prostate-specific antigen testing, lack sufficient specificity, particularly for African populations. Two GEO datasets, GSE6919 and GSE46602, were analyzed to identify differentially expressed genes (DEGs). Preprocessing and DEG identification were performed using GEO2R with thresholds of adjusted p-value < 0.05 and |log2 fold change| > 1.5. Intersection analysis revealed 223 commonly upregulated and 121 commonly downregulated genes across both datasets. Functional enrichment analysis using DAVID demonstrated significant enrichment in focal adhesion, extracellular matrix remodeling, TGF-β signaling, and other cancer-related pathways. Fourteen hub genes were selected for further validation. Among these, FGFR2, CCND2, NFE2L2, COL4A5, and COL4A6 showed significant dysregulation. Expression patterns, promoter methylation status, and overall survival relevance were validated using UALCAN based on TCGA data. These findings highlight key genes and pathways associated with prostate cancer progression and provide candidate biomarkers for further experimental validation and potential development of cost-effective diagnostic panels suitable for healthcare systems in Africa.
Keywords: Bioinformatics Analysis, Prostate cancer genomics, in silico biomarker discovery, Gene expression datasets.