
The GESMA dataset is a large-scale collection of real-world Ghanaian environmental soundscapes designed to support machine learning research in acoustic scene analysis and sound event classification, particularly in low-resource and underrepresented contexts. It comprises 22,193 uncompressed 44.1 kHz/16-bit WAV audio recordings (approximately 62.24 hours) captured across diverse urban, institutional, and community environments in Ghana, including markets, transport hubs, educational settings, public spaces, and human non-verbal acoustic contexts.
Audio recordings were collected in situ using consumer-grade smartphones under natural acoustic conditions to preserve authentic background noise, reverberation, and overlapping sound events. All files are organized using a hierarchical labeling structure (category → class → subclass) and are accompanied by a structured CSV metadata file containing contextual information such as location, environment type, time of day, and recording device. Annotations were manually verified through a multi-stage validation process to ensure label accuracy and consistency, and recordings containing intelligible speech or poor signal quality were excluded.
The dataset addresses a significant geographic and cultural gap in existing environmental audio resources, which are predominantly derived from Western contexts. GESMA can be reused for a wide range of applications, including environmental monitoring, sound event detection, accessibility and assistive technologies, smart-city research, and the development and benchmarking of context-aware machine listening models. Its use of consumer-grade recording devices also enables reproducibility and extension by other researchers.