Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

WildCLIP: Scene and animal attribute retrieval from camera trap data with domain-adapted vision-language models

Domaine:

environment and energygeospatial

Type de record:

datasetmodel
Créateur:
Gabeff, ValentinRußwurm, MarcTuia, DevisMathis, Alexander
Éditeur:
Zenodo
Hôte:avatar

#############

WildCLIP: Scene and animal attribute retrieval from camera trap data with domain-adapted vision-language models

#############

Authors: Valentin Gabeff, Marc Russwurm, Devis Tuia & Alexander Mathis

Affiliation: EPFL

Date: January, 2024

Link to the BiorXiv article: 

https://www.biorxiv.org/con…

--------------------------------

WildCLIP is a fine-tuned CLIP model that allows to retrieve camera-trap events with natural language from the Snapshot Serengeti dataset. This project intends to demonstrate how vision-language models may assist the annotation process of camera-trap datasets.

Here we provide the processed Snapshot Serengeti data used to train and evaluate WildCLIP, along with two versions of WildCLIP (model weights).

Details on how to run these models can be found in the project github repository.

Provided data (images and attribute annotations): 

The data consists of 380 x 380 image crops corresponding to the MegaDetector output of Snapshot Serengeti with a confidence threshold above 0.7. We considered only camera trap images containing single individuals.

A description of the original data can be found on LILA here, released under the Community Data License Agre….

We warmly thank the authors of LILA for making the MegaDetector outputs publicly available, as well as for structuring the dataset and facilitating its access.

Adapted CLIP model (model weights): 

WildCLIP models provided:

  • WildCLIP_vitb16_t1.pth: CLIP model with the ViT-B/16 visual backbone trained on data with captions following template 1.
  • WildCLIP_vitb16_t1t7_lwf.pth: CLIP model with the ViT-B/16 visual backbone trained on data with captions following templates 1 to 7, and with the additional VR-LwF loss.

We also provide the CSV files containing the train / val / test splits. The train / test splits follow camera split from LILA (lila.science). The validation split is custom, and also at the camera level.

  • train_dataset_crops_single_animal_template_captions_T1T7_ID.csv: Train set with captions from templates 1 through 7 (column "all captions") or template 1 only (column "template 1")
  • val_dataset_crops_single_animal_template_captions_T1T7_ID.csv: Validation set with captions from templates 1 through 7 (column "all captions") or template 1 only (column "template 1")
  • test_dataset_crops_single_animal_template_captions_T1T8T10.csv: Test set with captions from templates 1, 8, 9 and 10 (columns "all captions")

Details on how the models were generated can be found in the associated publication (preprint).

References: 

If you find our code, or weights, please cite:

@article{gabeff2023wildclip,
  title={WildCLIP: Scene and animal attribute retrieval from camera trap data with domain-adapted vision-language models},
  author={Gabeff, Valentin and Russwurm, Marc and Tuia, Devis and Mathis, Alexander},
  journal={bioRxiv},
  pages={2023--12},
  year={2023},
  publisher={Cold Spring Harbor Laboratory}
}

If you use the adapted Snapshot Serengeti data please also cite their article:

@article{swanson2015snapshot,
  title={Snapshot Serengeti, high-frequency annotated camera trap images of 40 mammalian species in an African savanna},
  author={Swanson, Alexandra and Kosmala, Margaret and Lintott, Chris and Simpson, Robert and Smith, Arfon and Packer, Craig},
  journal={Scientific data},
  volume={2},
  number={1},
  pages={1--14},
  year={2015},
  publisher={Nature Publishing Group}
}

Visit

doi.org

Tasks

computer visionimage-text retrieval

Licenses

info:eu-repo/semantics/openAccessCommunity Data License Agreement Permissive 1.0https://cdla.io/permissive-1-0

Similaires

Data release of Paying Attention to Other Animal Detections Improves Camera Trap Classification ModelsData and code release of Paying Attention to Other Animal Detections Improves Camera Trap Classification ModelsZero‐shot animal behaviour classification with vision‐language foundation modelsIntra-African Domain Shift in Wildlife Camera Trap AI]% {Intra-African Geographic Domain Shift in Wildlife Camera Trap Species Classification: A Comparative Study of Supervised and Zero-Shot Foundation ModelsData from: Identifying drivers of spatial variation in occupancy with limited replication camera trap dataGreenCrossingAI: A Camera Trap/Computer Vision Pipeline for Environmental Science Research Groups

Data release of Paying Attention to Other Animal Detections Improves Camera Trap Classification Models

Data release of: Paying Attention to Other Animal Detections Improves Camera Trap Classification Mod

Data and code release of Paying Attention to Other Animal Detections Improves Camera Trap Classification Models

Data and code release of: Paying Attention to Other Animal Detections Improves Camera Tra

Zero‐shot animal behaviour classification with vision‐language foundation models

International audience Understanding the behaviour of animals in their natural habita

Intra-African Domain Shift in Wildlife Camera Trap AI]% {Intra-African Geographic Domain Shift in Wildlife Camera Trap Species Classification: A Comparative Study of Supervised and Zero-Shot Foundation Models

This repository contains al

Data from: Identifying drivers of spatial variation in occupancy with limited replication camera trap data

Occupancy models are widely used in camera trap studies to analyze species presence, abundance, and

GreenCrossingAI: A Camera Trap/Computer Vision Pipeline for Environmental Science Research Groups

Camera traps have long been used by wildlife researchers to monitor and study animal behavior, popul