Logo Lanfrica

A Data-constrained Machine Learning Framework for District-level Crop Yield Prediction in Ghana

Domaine:

agriculture

Type de record:

paper
Créateur:
Pat
Éditeur:
Spr
Hôte:
Abstract Accurate crop yield prediction is essential for food security planning and agricultural policy formulation, particularly in developing economies where data availability is limited. Traditional yield forecasting methods based on historical averages and expert judgment are often insufficient to capture the complex, non-linear interactions that govern agricultural productivity. This study proposes a machine learning–based framework for district-level crop yield prediction in Ghana using long-term historical data obtained from the Ministry of Food and Agriculture (MOFA). A Random Forest regression model was developed and evaluated using a structured dataset spanning 2000–2024, comprising spatial, temporal, and agronomic variables including region, district, season, cultivated area, and production output. Model performance was benchmarked against baseline approaches such as Linear Regression, Decision Trees, and a Dummy Regressor using a realistic temporal train–validation–test split. Evaluation metrics included RMSE, MSE, and the coefficient of determination (R²). Results indicate that the Random Forest model significantly outperformed baseline models, achieving an RMSE of 0.82 tonnes/ha and an R² of 0.99 on unseen test data (2022–2024). Feature importance analysis revealed that temporal trends and spatial location were the most influential predictors of yield. The findings demonstrate that interpretable ensemble machine learning methods can deliver highly accurate and policy-relevant yield forecasts using structured historical data alone. The study provides a scalable and practical framework for data-driven agricultural planning in Ghana and similar contexts.