Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

HumBugDB: A Large-scale Acoustic Mosquito Dataset

Domaine:

healthcarenatural language processing

Type de record:

datasetpaper
Créateur:
KisSinka, MarianneCobRaf
Éditeur:
arXiv
Hôte:avatar
This paper presents the first large-scale multi-species dataset of acoustic recordings of mosquitoes tracked continuously in free flight. We present 20 hours of audio recordings that we have expertly labelled and tagged precisely in time. Significantly, 18 hours of recordings contain annotations from 36 different species. Mosquitoes are well-known carriers of diseases such as malaria, dengue and yellow fever. Collecting this dataset is motivated by the need to assist applications which utilise mosquito acoustics to conduct surveys to help predict outbreaks and inform intervention policy. The task of detecting mosquitoes from the sound of their wingbeats is challenging due to the difficulty in collecting recordings from realistic scenarios. To address this, as part of the HumBug project, we conducted global experiments to record mosquitoes ranging from those bred in culture cages to mosquitoes captured in the wild. Consequently, the audio recordings vary in signal-to-noise ratio and contain a broad range of indoor and outdoor background environments from Tanzania, Thailand, Kenya, the USA and the UK. In this paper we describe in detail how we collected, labelled and curated the data. The data is provided from a PostgreSQL database, which contains important metadata such as the capture method, age, feeding status and gender of the mosquitoes. Additionally, we provide code to extract features and train Bayesian convolutional neural networks for two key tasks: the identification of mosquitoes from their corresponding background environments, and the classification of detected mosquitoes into species. Our extensive dataset is both challenging to machine learning researchers focusing on acoustic identification, and critical to entomologists, geo-spatial modellers and other domain experts to understand mosquito behaviour, model their distribution, and manage the threat they pose to humans. Accepted at the 35th Conference on Neural Information Processing Systems (NeurIPS 2021) Track on Datasets and Benchmarks. 10 pages main, 39 pages including appendix. This paper accompanies the dataset found at zenodo.org with corresponding code at github.com

Visit

doi.orgarxiv.org

Tasks

speech processing

Tags

Sound (cs.SD)Computer Vision and Pattern Recognition (cs.CV)Audio and Speech Processing (eess.AS)FOS: Computer and information sciencesFOS: Computer and information sciencesFOS: Electrical engineering, electronic engineering, information engineeringFOS: Electrical engineering, electronic engineering, information engineeringE.0; I.2.1; J.3

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

Large-scale automatic acoustic monitoring of African forest elephants' calls in the terrestrial acoustic recordingsA subset of large-scale EEG dataset (India + Tanzania)Fidel: A Large-Scale Sentence Level Amharic OCR DatasetFraud Detection Using Large-scale Imbalance DatasetMassiveSumm: a very large-scale, very multilingual, news summarisation datasetCAT2000: A Large Scale Fixation Dataset for Boosting Saliency Research

Large-scale automatic acoustic monitoring of African forest elephants' calls in the terrestrial acoustic recordings

African forest elephants live in the rain forests of western and central Africa. The dense habitat p

A subset of large-scale EEG dataset (India + Tanzania)

Dataset ID: ds007358 Vianney2026 Canonical aliases: Vianney2025 At a glance: EEG · Resting State re

Fidel: A Large-Scale Sentence Level Amharic OCR Dataset

Abstract The Ethiopic script used in the Amharic Language presents persistent chal

Fraud Detection Using Large-scale Imbalance Dataset

In the context of machine learning, an imbalanced classification problem states to a dataset in whic

MassiveSumm: a very large-scale, very multilingual, news summarisation dataset

Current research in automatic summarisation is unapologetically anglo-centered–a persistent state-of-affairs, which also predates neural net approaches. High-quality automatic summarisation datasets are notoriously expensive to create, posing a challenge for any la

CAT2000: A Large Scale Fixation Dataset for Boosting Saliency Research

Saliency modeling has been an active research area in computer vision for about two decades. Existin