Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Crop Recommendation Dataset with SMOTE, VAE and CropGAN Synthetic Generated Dataset Variants

Domain:

agriculture

Record type:

dataset
Creator:
KwaEkeAjo
Publisher:
Zenodo
Host:avatar

Crop Recommendation Dataset with Synthetic Data Augmentation

1. Overview

This dataset was developed to support research in data-driven crop recommendation systems. It contains soil and environmental attributes used to predict suitable crops for cultivation. In addition to the original dataset, this repository includes preprocessed data and synthetically generated samples produced using advanced data augmentation techniques.

The dataset is designed to address class imbalance and improve model performance in precision agriculture applications.

2. Dataset Contents

The repository includes the following files:

  1. Original_dataset.csv
    Original dataset collected from field survey
  2. Original_preprocessed_dataset.csv
    Cleaned and transformed dataset used for model training (includes encoding and feature engineering)
  3. Smote_dataset.csv
    Synthetic dataset generated using SMOTE to address class imbalance
  4. Vae_dataset.csv
    Synthetic dataset generated using Variational Autoencoder (VAE)
  5. Cropgan_dataset.csv
    Synthetic dataset generated using the proposed GAN-based framework (CropGAN)
  6. Data_dictionary.csv
    Description of all features and their meanings

3. Features Description

Typical attributes include:

  1. Soil properties (e.g., pH, soil_texture, nutrients (N, P, K))
  2. Environmental factors (e.g., rainfall, temperature, humidity)
  3. Crop label (target variable)

Categorical variables such as soil texture were encoded during preprocessing.

4. Data Collection

The dataset used in this study was originally collected by Kwaghtyo et al. [6]. The initial dataset consisted of 2,200 samples. To enhance model robustness and enable comprehensive experimentation, the dataset was expanded to 5,000 samples.

The original 2,200 samples were retained to ensure comparability with previous studies, while an additional 2,800 samples were collected using a random sampling approach across farm fields in Yandev District, Gboko Local Government Area (LGA), Benue State, Nigeria.

Data collection for the additional 2,800 samples was conducted across two farming seasons (2024 and 2025) to capture seasonal variability. Soil samples were collected at a depth of 0-30 cm using standard soil auger techniques and analysed in the laboratory to determine soil nutrient composition and physicochemical properties.

Climatic variables were also recorded, with temperature ranging from 25.0 °C to 33.5 °C and rainfall ranging from 900 mm to 1,200 mm.

The dataset comprises nine attributes: nitrogen (N), phosphorus (P), potassium (K), soil pH, humidity, temperature, rainfall, soil texture, and crop label. The target crops include maize, rice, pepper, soybean, beans, orange, guinea corn, cassava, tomatoes, and yam, which are commonly cultivated in the study area.

5. Preprocessing Steps

The following preprocessing steps were applied:

  1. Handling missing values
  2. Normalisation/standardisation
  3. One-hot encoding of categorical variables
  4. Feature selection (where applicable)

6. Synthetic Data Generation

To mitigate class imbalance, the following techniques were applied:

  1. SMOTE (Synthetic Minority Oversampling Technique)
  2. Variational Autoencoder (VAE)
  3. Proposed GAN-based model (CropGAN)

Each method generates additional samples while preserving the statistical properties of the original dataset.

7. Intended Use

This dataset is intended for:

  1. Crop recommendation research
  2. Machine learning model development
  3. Class imbalance studies
  4. Synthetic data evaluation

8. Limitations

  1. The dataset may not generalise to all geographic regions
  2. Synthetic data may introduce minor distributional deviations
  3. Results depend on preprocessing and model configuration

9. Reproducibility

All datasets used in the study (original, original_preprocessed, and the synthetic variants) are provided to ensure reproducibility of experimental results.

10. Citation

If you use this dataset, please cite:

Kwaghtyo, K.D., Eke, C.I. and Ajon, A.T. (2026). Crop Recommendation Dataset with Synthetic Data Augmentation.

11. Contact

For questions or collaboration:
Dekera Kenneth Kwaghtyo
Federal University of Lafia, PMB 146, Nasarawa State, Nigeria
dekera.kwaghtyo@student.fulafia.edu.ng

Visit

doi.org

Tasks

text classification

Languages

Deg

Tags

crop recommendationSynthetic data generationCropGANSMOTEVAE

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Crop Recommendation using Soil Properties and Weather Prediction DatasetTelecom Recommendation Dataset (Algeria) simulatedGenerated dataset for VH band.Deepfake-Synthetic-20K DatasetAfrican Cerebral Palsy Synthetic DatasetSomali ASR Synthetic YouTube Dataset

Crop Recommendation using Soil Properties and Weather Prediction Dataset

The dataset is comprehensive, encompassing various key factors critical to machine learning-based cr

Telecom Recommendation Dataset (Algeria) simulated

Predict churn and recommend mobile plans using data (djezzy, ooreedoo, mobilis)

Generated dataset for VH band.

The identification of various objects and species found in nature is of great importance tod

Deepfake-Synthetic-20K Dataset

The Deepfake-Synthetic-20K dataset is a significant contribution to the field of digital forensics a

African Cerebral Palsy Synthetic Dataset

⚠️ Synthetic dataset — Parameterized from published SSA literature, not real observations. Not suita

Somali ASR Synthetic YouTube Dataset

A Somali-language speech dataset derived from YouTube audio, intended for training and evaluating au