Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Detection and classification of vocal productions in large scale audio recordings

Domaine:

environment and energy

Type de record:

paperdataset
Créateur:
BonPudFreLeg
Éditeur:
InsLabInsLab
Éditeur:
CCSD
Hôte:avatar
We propose an automatic data processing pipeline to extract vocal productions from large-scale natural audio recordings and classify these vocal productions. The pipeline is based on a deep neural network and adresses both issues simultaneously. Though a series of computationel steps (windowing, creation of a noise class, data augmentation, re-sampling, transfer learning, Bayesian optimisation), it automatically trains a neural network without requiring a large sample of labeled data and important computing resources. Our end-to-end methodology can handle noisy recordings made under different recording conditions. We test it on two different natural audio data sets, one from a group of Guinea baboons recorded from a primate research center and one from human babies recorded at home. The pipeline trains a model on 72 and 77 minutes of labeled audio recordings, with an accuracy of 94.58% and 99.76%. It is then used to process 443 and 174 hours of natural continuous recordings and it creates two new databases of 38.8 and 35.2 hours, respectively. We discuss the strengths and limitations of this approach that can be applied to any massive audio recording.

Visit

doi.org

Tasks

speech processing

Tags

detectionclassificationneural networktransfer learningvocalization[STAT]Statistics [stat][INFO.INFO-SD]Computer Science [cs]/Sound [cs.SD][SCCO]Cognitive science[STAT.AP]Statistics [stat]/Applications [stat.AP][STAT.ML]Statistics [stat]/Machine Learning [stat.ML]

Licenses

info:eu-repo/semantics/OpenAccess

Similaires

Fraud Detection Using Large-scale Imbalance DatasetDetection of Large-Scale Floods Using Google Earth Engine and Google Colab[Pawnee audio recordings][Pomo audio recordings]Large-scale automatic acoustic monitoring of African forest elephants' calls in the terrestrial acoustic recordingsClassification of time series of Sentinel-2 images for large scale mapping in Cameroon

Fraud Detection Using Large-scale Imbalance Dataset

In the context of machine learning, an imbalanced classification problem states to a dataset in whic

Detection of Large-Scale Floods Using Google Earth Engine and Google Colab

International audience This paper presents an operational approach for detecting floo

[Pawnee audio recordings]

http://cla.berkeley.edu/item/1465

[Pomo audio recordings]

http://cla.berkeley.edu/item/1009

Large-scale automatic acoustic monitoring of African forest elephants' calls in the terrestrial acoustic recordings

African forest elephants live in the rain forests of western and central Africa. The dense habitat p

Classification of time series of Sentinel-2 images for large scale mapping in Cameroon

International audience Sentinel-2 satellites provide dense image time series exhibiti