
Crop Recommendation Dataset with Synthetic Data Augmentation
1. Overview
This dataset was developed to support research in data-driven crop recommendation systems. It contains soil and environmental attributes used to predict suitable crops for cultivation. In addition to the original dataset, this repository includes preprocessed data and synthetically generated samples produced using advanced data augmentation techniques.
The dataset is designed to address class imbalance and improve model performance in precision agriculture applications.
2. Dataset Contents
The repository includes the following files:
3. Features Description
Typical attributes include:
Categorical variables such as soil texture were encoded during preprocessing.
4. Data Collection
The dataset used in this study was originally collected by Kwaghtyo et al. [6]. The initial dataset consisted of 2,200 samples. To enhance model robustness and enable comprehensive experimentation, the dataset was expanded to 5,000 samples.
The original 2,200 samples were retained to ensure comparability with previous studies, while an additional 2,800 samples were collected using a random sampling approach across farm fields in Yandev District, Gboko Local Government Area (LGA), Benue State, Nigeria.
Data collection for the additional 2,800 samples was conducted across two farming seasons (2024 and 2025) to capture seasonal variability. Soil samples were collected at a depth of 0-30 cm using standard soil auger techniques and analysed in the laboratory to determine soil nutrient composition and physicochemical properties.
Climatic variables were also recorded, with temperature ranging from 25.0 °C to 33.5 °C and rainfall ranging from 900 mm to 1,200 mm.
The dataset comprises nine attributes: nitrogen (N), phosphorus (P), potassium (K), soil pH, humidity, temperature, rainfall, soil texture, and crop label. The target crops include maize, rice, pepper, soybean, beans, orange, guinea corn, cassava, tomatoes, and yam, which are commonly cultivated in the study area.
5. Preprocessing Steps
The following preprocessing steps were applied:
6. Synthetic Data Generation
To mitigate class imbalance, the following techniques were applied:
Each method generates additional samples while preserving the statistical properties of the original dataset.
7. Intended Use
This dataset is intended for:
8. Limitations
9. Reproducibility
All datasets used in the study (original, original_preprocessed, and the synthetic variants) are provided to ensure reproducibility of experimental results.
10. Citation
If you use this dataset, please cite:
Kwaghtyo, K.D., Eke, C.I. and Ajon, A.T. (2026). Crop Recommendation Dataset with Synthetic Data Augmentation.
11. Contact
For questions or collaboration:
Dekera Kenneth Kwaghtyo
Federal University of Lafia, PMB 146, Nasarawa State, Nigeria
dekera.kwaghtyo@student.fulafia.edu.ng