SUMMARY
Machine learning models offer powerful predictive capabilities for geoscientific applications but remain limited by their ‘black-box’ nature and lack of rigorous uncertainty quantification. We developed a comprehensive, generalizable uncertainty quantification framework that decomposes predictive uncertainty into aleatoric and epistemic components using Quantile Regression Forests. Additionally, we applied unsupervised k-means clustering to isolate homogeneous data regimes, thereby reducing aleatoric uncertainty across spatially heterogeneous geoscientific data sets. To facilitate interpretation and quality assessment, we introduced five spatial diagnostic tools: bandwidth, variance, robustness, confidence and explainability maps that characterize prediction reliability and identify dominant uncertainty sources. To demonstrate the framework’s applicability, we tested it on three synthetic data sets varying in size and a real-world geothermal heat flow application with 14 geophysical observables across continental Africa. Results show that clustering substantially reduces aleatoric uncertainty while maintaining stable epistemic uncertainty. Clustering also improves predictive accuracy and sharpens prediction intervals, with gains most pronounced in homogeneous regions. Applied to the African geothermal heat flow, the framework reveals region-specific geological controls (lithospheric architecture dominates stable cratons, while tectonic proximity governs active rift zones) and guides targeted data collection by distinguishing high-epistemic regions requiring additional sampling from high-aleatoric zones needing improved observables. While theoretically applicable to other geographic regions and geophysical data sets, the framework’s performance in different geological settings requires validation. This interpretable, uncertainty-aware approach enhances trustworthiness of predictions in spatially heterogeneous, data-sparse geoscientific problems.