Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Great Ape Detection in Challenging Jungle Camera Trap Footage via Attention-Based Spatial and Temporal Feature Blending

Domain:

environment and energygeospatial

Record type:

papermodeldataset
Creator:
YanMirBur
Publisher:
arXiv
Host:avatar
We propose the first multi-frame video object detection framework trained to detect great apes. It is applicable to challenging camera trap footage in complex jungle environments and extends a traditional feature pyramid architecture by adding self-attention driven feature blending in both the spatial as well as the temporal domain. We demonstrate that this extension can detect distinctive species appearance and motion signatures despite significant partial occlusion. We evaluate the framework using 500 camera trap videos of great apes from the Pan African Programme containing 180K frames, which we manually annotated with accurate per-frame animal bounding boxes. These clips contain significant partial occlusions, challenging lighting, dynamic backgrounds, and natural camouflage effects. We show that our approach performs highly robustly and significantly outperforms frame-based detectors. We also perform detailed ablation studies and validation on the full ILSVRC 2015 VID data corpus to demonstrate wider applicability at adequate performance levels. We conclude that the framework is ready to assist human camera trap inspection efforts. We publish code, weights, and ground truth annotations with this paper. Accepted by ICCV workshop 2019

Visit

doi.orgarxiv.org

Tasks

computer vision

Tags

Computer Vision and Pattern Recognition (cs.CV)FOS: Computer and information sciencesFOS: Computer and information sciences

Licenses

arXiv.org perpetual, non-exclusive licensehttp://arxiv.org/licenses/nonexclusive-distrib/1.0/

Similar

Attention-Based LSTM for Sign Language Recognition Leveraging Spatial-Temporal Keypointcamera-trap-datasetData release of Paying Attention to Other Animal Detections Improves Camera Trap Classification ModelsGreat ape autosomal STR sequence dataData and code release of Paying Attention to Other Animal Detections Improves Camera Trap Classification ModelsThe Great Ape Dictionary Video Data Ark

Attention-Based LSTM for Sign Language Recognition Leveraging Spatial-Temporal Keypoint

Sign language is a crucial means of communication for the Deaf and hard-of-hearing communities. Most

camera-trap-dataset

This dataset has camera trap images of wildlife species from a conservancy in Kenya. They are based

Data release of Paying Attention to Other Animal Detections Improves Camera Trap Classification Models

Data release of: Paying Attention to Other Animal Detections Improves Camera Trap Classification Mod

Great ape autosomal STR sequence data

DNA sequences of the orthologs of human forensic autosomal STR loci in chimpanzees, bonobos and c

Data and code release of Paying Attention to Other Animal Detections Improves Camera Trap Classification Models

Data and code release of: Paying Attention to Other Animal Detections Improves Camera Tra

The Great Ape Dictionary Video Data Ark

We study the behaviour and cognition of wild apes and other species (elephants, corvids, do