This dataset was developed as part of the Birds’ Detector and Repellent System for Large-Scale Smart Farming project to support research in smart agriculture, computer vision, acoustic sensing, TinyML, embedded artificial intelligence (AI), and Internet of Things (IoT)-based crop protection systems. The dataset was collected from rice farming environments and nearby park ecosystems in Kenya and Rwanda using distributed sensing devices including 64MP Raspberry Pi Arducam OwlSight cameras, AudioMoth recorders, MEMS microphones, and embedded edge computing systems.
The dataset consists of multimodal environmental data containing image and acoustic recordings collected under real agricultural field conditions. A total of over 700 environmental images and motion-based image sequences were captured using deployed cameras positioned across rice farms and surrounding environments. The image dataset was designed primarily for motion-aware bird detection, environmental monitoring, and bird-versus-background classification rather than species-level image annotation. The images include environmental motion sequences, vegetation-rich scenes, sky-background imagery, long-range agricultural views, and dynamic field conditions captured under varying illumination and weather conditions.
The acoustic dataset contains over 20,000 audio clips of approximately 2–3 seconds duration collected over six months using AudioMoth devices and MEMS-based acoustic sensors deployed within rice farming areas. The recordings were processed using Audacity software and annotated using the BirdNET platform into six classes: common waxbill, red-billed quelea, village weaver, yellow-fronted canary, other birds, and environmental noise. The environmental noise category includes wind interference, rainfall, insect sounds, human activities, and farm equipment noise collected under natural field conditions.
Image preprocessing involved resizing, normalization, frame differencing, adaptive thresholding, contour extraction, and bounding box generation. Acoustic preprocessing included segmentation, noise filtering, spectrogram generation, and Mel-Frequency Cepstral Coefficients (MFCCs) extraction for acoustic feature representation.
The dataset supports research in precision agriculture, environmental monitoring, edge AI, multimodal machine learning, acoustic classification, embedded AI systems, and intelligent crop protection technologies operating under resource-constrained deployment environments.