This research evaluates machine learning approaches for gold prospectivity mapping in Bindura District, Zimbabwe, under severe data constraints typical of underexplored mineral terrains. The study addresses three research questions: (1) how machine learning algorithms predict gold prospectivity using integrated geological, geochemical, remote sensing, and structural data; (2) which algorithms achieve optimal performance with limited training data; and (3) how variable combinations influence prediction accuracy.
Three algorithmically distinct methods were compared under identical conditions: One-Class Support Vector Machine (SVM), Isolation Forest, and XGBoost. Using only four verified gold deposits as training examples and fourteen integrated predictor variables, each algorithm generated prospectivity predictions across 51,134 pixels (approximately 46 km²) in mid-Bindura District. The experimental design enabled systematic performance comparison through standardized data preprocessing, model training, and output generation workflows implemented in three independent Python scripts.
Results demonstrated that Isolation Forest achieved superior performance, successfully identifying 6.9 km² of high-priority exploration targets (14.98% of study area) with clear spatial discrimination between prospective and non-prospective zones. In divergence, One-Class SVM completely failed, making uniform probability assignments (range: 0.4622-0.4705) with zero spatial discrimination. XGBoost created overly conservative predictions (mean probability: 0.0705) that acknowledged no definitive high significance targets despite producing 3,413 unique probability values.
A key finding revealed that algorithm performance depended strongly on data quality and quantity rather than inherent algorithmic capability. The experimental failures and limitations principally reflected data insufficiency only four training deposits, sparse geochemical sampling demanding extensive interpolation, and incomplete spatial coverage rather than fundamental algorithmic limitations. All three methods would likely prove substantially better quality and potentially comparable performance with expanded training datasets (30-50 deposits), systematic geochemical grid sampling, and complete district coverage.
The study establishes that machine learning based prospectivity mapping is achievable in data scarce circumstances but functions at reduced effectiveness. The supported exploration targets offer actionable guidance for field campaigns, though field validation is essential. Future work ought to prioritize intensifying training datasets and refining geochemical sampling density to realise the full potential of machine learning methods for mineral exploration funding in Zimbabwe's greenstone belts.