Globally, food production relies heavily on a few major staple crops, including maize, wheat, and rice. While these crops are highly productive, their dominance makes agricultural systems vulnerable to climate-related challenges such as prolonged droughts, intense floods, and increasingly unpredictable rainfall. This vulnerability highlights the need for crop diversification as a strategy to enhance climate resilience, enabling agricultural systems to maintain productivity under variable environmental conditions and support long-term sustainability. Neglected and underutilised crop species (NUS) such as taro (Colocasia esculenta) offer a promising pathway for achieving agricultural diversification, food and nutrition security, and climate resilience. Taro’s adaptability to marginal environments, low input requirements, and high nutritional value make this NUS suitable for sustaining food production and livelihoods under increasing climate variability. Despite this potential, NUS remain marginalised in research and agricultural planning. A major barrier is the lack of credible, consolidated, and spatially explicit data describing their agronomic requirements, yield potential, and water use across diverse agroecological zones. This persistent data gap constrains the calibration and validation of crop simulation models (CSMs), which are essential for evaluating crop performance, assessing land suitability, and supporting agricultural planning and policy development. To address this gap, this study developed a comprehensive methodological framework to improve the availability and usability of data for taro. Specifically, the research aimed to: (i) generate reliable, high-quality modelled datasets for taro using a calibrated CSM; (ii) use these datasets to develop a national-scale taro database, and conduct spatial assessments of the crop’s performance, including cycle length, potential yield, and both crop and nutritional water productivity; and (iii) use the database to develop data-driven yield prediction models capable of providing simplified yield estimates.
A systematic scoping review and bibliometric analysis (Chapter 2) were undertaken to establish the global status of agricultural databases for crop modelling. The review confirmed that research and database development are largely focused on major staple crops (i.e. maize, soybean, and wheat), with a clear geographical bias towards the Global North, with NUS largely excluded. The review revealed a profound data and knowledge gap, characterised by insufficient data for NUS and limited availability of measured soil data needed for robust model calibration and validation. It further identified that using modelled data is a viable strategy to bridge the data gap but emphasised that the credibility of this approach depends entirely on the availability of a well-calibrated and validated CSM. The second objective (Chapter 3) was therefore to recalibrate and validate the AquaCrop model for selected NUS (i.e. sweet potato and taro) using multi-location experimental datasets and an iterative parameter refinement process. Key adjustments were made to crop parameters governing phenology (where physiological maturity was based on tuber mass stabilisation rather than senescence), canopy development, and the crop's response to water stress. The recalibration was successful, achieving strong agreements for optimal conditions between simulated and observed data for canopy cover (R2 up to 0.954), biomass (NSE up to 0.975), and final yield (absolute deviations ≤ 6%). Validation across three locations confirmed the model’s reliability in predicting final yield under non-stressed conditions, although its accuracy declined under water-limited environments. Importantly, this chapter also demonstrated the importance of running AquaCrop in growing-degree-day (GDD) mode, particularly when validating the model against datasets from sites with climatic conditions that differ from the calibration environment. Using GDD-based phenological growth improved the accuracy of the model at other locations and strengthened its suitability for the subsequent national-scale assessment.
In Chapter 4, AquaCrop was applied in GDD mode to conduct a national-scale spatial assessment of taro’s potential, fulfilling the study’s main aim of developing a comprehensive database for the crop. The model was run for 5838 relatively homogenous altitude zones across southern Africa (including South Africa, Eswatini, and Lesotho), with 50 years of historical climate data input that produced up to 49 consecutive seasons of simulated data for five planting dates. A key advancement was the inclusion of a frost occurrence that has been omitted from previous taro spatial modelling efforts. This was achieved by integrating a frost risk threshold (5.4°C) reported in the literature, below which crop growth immediately ceases. This enabled the generation of more realistic maps of crop cycle duration, potential yield, and both crop and nutritional water productivity. The findings identified optimal planting dates and quantified the impact of frost, which rendered inland regions at high-altitude as unsuitable for taro production. An economic analysis, based on actual farm-level costs and market prices, demonstrated taro's economic viability, with potential net profits of approximately R 38 000 ha-1 and an economic water productivity of approximately R 11 m-3. These findings highlighted taro’s potential as a profitable and climate-resilient crop for small-scale and emerging farmers. The process-based modelling in Chapter 4 revealed a practical limitation of CSMs, which are computationally expensive and technically complex to run, making them impractical for many end-users to apply. To bridge this gap, the extensive crop database provided training and testing datasets to facilitate the development of data-driven yield models (Chapter 5). A generalised linear model (GLM) and a random forest model were developed to predict taro yield using minimal inputs (i.e. rainfall and reference evapotranspiration accumulated over the growing season, and total available water). Model performance varied spatially, with some agroecological zones showing a closer agreement with AquaCrop simulations than others. The GLM provided coefficients that enabled the development of a yield prediction equation. Validation against observed yields highlighted the potential of the GLM as a transparent, data-efficient tool for on-farm decision support.
While this study establishes a robust and transferable framework, further work is required to strengthen and expand its applicability. The recalibrated AquaCrop parameters for taro should be validated using additional, independent, multi-location experimental datasets to confirm their reliability across diverse environments. Similarly, the data-driven yield models need to be evaluated against broader yield observations to critically assess their predictive accuracy. Future simulations incorporating projected climate scenarios are also essential to assess taro’s yield resilience to the changing climate, thereby generating vital insights for regional adaptation planning. Refinement of the data-driven yield models through the inclusion of additional predictor variables and exploration of advanced ensemble approaches could further improve predictive performance and generalisability. Finally, the complete methodological framework developed in this study should be extended to other high-priority NUS such as bambara nut, cowpea, and sweet potato to support the creation of a comprehensive, multi-crop database that advances agricultural diversification and climate-resilient food system development.
In conclusion, the study addressed the stated objectives by developing a comprehensive taro database, applying it to conduct spatial assessments of crop performance, and using it to develop simplified data-driven yield prediction models. The marginalisation of NUS, particularly taro, is perpetuated by a critical data deficit and systemic issues such as limited investment in targeted research and policy constraints that hinder the integration of NUS into conventional agricultural systems. The recalibration and validation of the AquaCrop model provided a credible foundation for generating the simulated data required to construct the taro database. Application of the model at a national scale produced a spatially explicit database describing taro’s crop cycle duration, potential yield, and both crop and nutritional water productivity across southern Africa. The database further enabled the development of simplified data-driven yield prediction models capable of providing accessible yield estimates using limited input variables. By developing a complete methodological framework from model recalibration and validation to creating a comprehensive database of simulated data, which enabled spatial analyses of crop performance and the development of data-driven yield prediction models, this study provides a comprehensive assessment of taro’s agronomic and nutritional potential. The findings demonstrate how crop simulation modelling combined with data-driven approaches can improve the availability of credible agronomic information for data-scarce crops. This work establishes a framework for enhancing data availability for other data-scarce NUS, offering a viable pathway to enhance climate-smart agricultural diversification and improve food and nutrition security under increasingly variable climatic conditions.