Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Fusion Complexity Inversion: Why Simpler Cross View Modules Outperform SSMs and Cross View Attention Transformers for Pasture Biomass Regression

Domaine:

agriculture

Type de record:

paper
Créateur:
Man
Hôte:avatar
Accurate estimation of pasture biomass from agricultural imagery is critical for sustainable livestock management, yet existing methods are limited by the small, imbalanced, and sparsely annotated datasets typical of real world monitoring. In this study, adaptation of vision foundation models to agricultural regression is systematically evaluated on the CSIRO Pasture Biomass benchmark, a 357 image dual view dataset with laboratory validated, component wise ground truth for five biomass targets, through 17 configurations spanning four backbones (EfficientNet-B3 to DINOv3-ViT-L), five cross view fusion mechanisms, and a 4x2 metadata factorial. A counterintuitive principle, termed "fusion complexity inversion", is uncovered: on scarce agricultural data, a two layer gated depthwise convolution (R^2 = 0.903) outperforms cross view attention transformers (0.833), bidirectional SSMs (0.819), and full Mamba (0.793, below the no fusion baseline). Backbone pretraining scale is found to monotonically dominate all architectural choices, with the DINOv2 -> DINOv3 upgrade alone yielding +5.0 R^2 points. Training only metadata (species, state, and NDVI) is shown to create a universal ceiling at R^2 ~ 0.829, collapsing an 8.4 point fusion spread to 0.1 points. Actionable guidelines for sparse agricultural benchmarks are established: backbone quality should be prioritized over fusion complexity, local modules preferred over global alternatives, and features unavailable at inference excluded. Accepted to CVPR: Vision for Agriculture Workshop 2026 (Withdrawn)

Visit

arxiv.org

Tasks

computer vision

Languages

VunjoZimba

Tags

Computer Vision and Pattern RecognitionMachine Learning

Similaires

CLIF-Net: Intersection-guided Cross-view Fusion Network for Infection Detection from Cranial UltrasoundToken-Region Guided Cross-Attention Fusion for Multimodal Affect InterpretationGeo-R1: Unlocking VLM Geospatial Reasoning with Cross-View Reinforcement LearningMulMoSenT: Multimodal Sentiment Analysis for a Low-Resource Language Using Textual-Visual Cross-Attention and FusionOrganisational politics and its influence on employee engagement A cross-cultural view of Austria and NigeriaA Pilot Study of World View of Black and White South African Adolescent Pupils: Implications for Cross-Cultural Counselling

CLIF-Net: Intersection-guided Cross-view Fusion Network for Infection Detection from Cranial Ultrasound

Abstract This paper addresses the problem of detecting possible serious bacterial

Token-Region Guided Cross-Attention Fusion for Multimodal Affect Interpretation

Automated analysis of multimodal content on social networks has become a critical task for understan

Geo-R1: Unlocking VLM Geospatial Reasoning with Cross-View Reinforcement Learning

We introduce Geo-R1, a reasoning-centric post-training framework that unlocks geospatial reasoning i

MulMoSenT: Multimodal Sentiment Analysis for a Low-Resource Language Using Textual-Visual Cross-Attention and Fusion

First-ever Bengali Multimodal Sentiment Analysis (BMSA) corpus and details in https://www.sciencedir

Organisational politics and its influence on employee engagement A cross-cultural view of Austria and Nigeria

This research investigates the connection that exists between perceived organisational powerplay and

A Pilot Study of World View of Black and White South African Adolescent Pupils: Implications for Cross-Cultural Counselling

Research on cross-cultural counselling and psychotherapy began to receive emphasis in the 1970s in t