Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Sample Earth: Machine-Learning–Ready Land-Cover Reference Dataset

Domain:

agriculturegeospatial

Record type:

dataset
Creator:
VanLuoPerTel
Editor:
BioBioAll
Publisher:
Har
Host:avatar
This dataset is part of the Sample Earth initiative, a global effort to build open, high-quality reference data for improving the accuracy and inclusiveness of land-cover maps. It contains GPS-located land-cover samples that can be used to train and validate AI models that generate detailed, accurate maps, with a focus on coffee and cocoa production systems.
The data were collected across Vietnam and Ghana, combining expert interpretation of high-resolution satellite imagery (Google Earth, Planet) with a smaller subset of ground-truth observations. Each point is labeled and quality-controlled to represent a diverse range of land-cover types commonly found within and around smallholder production areas. The classification scheme includes 10 main classes (such as coffee, cocoa, orchard, natural forests) and 68 sub-classes (such as full sun coffee, coffee intercropped with black pepper).
While the primary goal is to distinguish coffee and cocoa systems from other land uses, the dataset also supports broader applications such as agricultural monitoring, deforestation analysis, ecosystem service mapping, land-use planning, and suitability modeling.
By providing transparent, well-validated training data, this dataset contributes to Sample Earth’s broader objective: strengthening AI-based land monitoring tools and supporting global efforts, including the EU Deforestation Regulation (EUDR), to ensure sustainable, deforestation-free agricultural supply chains.
The dataset is designed to grow continuously, incorporating new commodities, timeframes, and countries over time.
The data is released under a Creative Commons Attribution–NonCommercial license. Entities wishing to use the data for commercial purposes are encouraged to contact us to establish a tailored data-sharing agreement.

Methodology:The dataset was developed primarily through expert visual interpretation of high-resolution satellite imagery from Google Earth and Planet, collected between 2019 and 2022. A smaller subset of points in the Central Highlands of Vietnam was derived from field observations, providing additional ground-truth validation.
To enhance interpreter accuracy and contextual understanding, field visits and Google Street View assessments were conducted in both Vietnam and Ghana. These activities helped experts better recognize local land-use patterns and distinguish among different crop and landscape types.
All sample points were digitized and standardized using QGIS, with attributes including class ID, crop type, sampling date, and associated metadata to ensure consistency and interoperability.
This combined approach of expert interpretation, localized training, and structured data management ensured a high-quality, consistent, and machine-learning–ready dataset suitable for land-cover mapping and model training workflows. High resolution imagery for interpretationHigh resolution imagery for interpretationHigh resolution imagery for interpretationHigh resolution imagery for interpretationHigh resolution imagery for interpretationHigh resolution imagery for interpretationHigh resolution imagery for interpretationHigh resolution imagery for interpretation

Visit

doi.orgdataverse.harvard.edu

Tasks

computer visionimage classification

Tags

Agricultural SciencesEarth and Environmental Sciencesreference pointscommoditiescoffeecocoageospatial dataobservational datareference dataremote sensing+10

Licenses

info:eu-repo/semantics/openAccessCustom terms specific to this datasethttps://dataverse.harvard.edu/api/datasets/:persistentId/versions/1.1/customlicense?persistentId=doi:10.7910/DVN/U7HWY1

Similar

Sample Earth — Tree Crop Reference Dataset — (Colombia, Ghana)Digital Earth Africa - Land CoverDigital Earth Africa - Land Cover DataLandsat 8Bands’ 1 to 7 spectral vectors plus machine learning to improve land use/cover classification using Google Earth EngineMapping Maize Cropland and Land Cover in Semi-Arid Region in Northern Nigeria Using Machine Learning and Google Earth EngineLand Cover Reference Data Series, OBSYDYA project, Benin

Sample Earth — Tree Crop Reference Dataset — (Colombia, Ghana)

This dataset contains GPS-located land-cover samples that can be used to train and validate AI model

Digital Earth Africa - Land Cover

Environment monitoring, urban growth, Land-use analysis, climate adaptation Notes / challenges: Req

Digital Earth Africa - Land Cover Data

Environmental monitoring, urban growth Notes / challenges: Requires large storage/processing capaci

Landsat 8Bands’ 1 to 7 spectral vectors plus machine learning to improve land use/cover classification using Google Earth Engine

This paper explores a spectral vector-based methodology on Landsat 8 bands of the visible wavelength

Mapping Maize Cropland and Land Cover in Semi-Arid Region in Northern Nigeria Using Machine Learning and Google Earth Engine

The monitoring of crop quantity and quality is vital for global food security. National food securit

Land Cover Reference Data Series, OBSYDYA project, Benin

The OBSYDYA project (Observatoire Pilote des Paysages et des Dynamiques Agricoles) aims to establish