Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

AgriGPT-VL: Agricultural Vision-Language Understanding Suite

Domaine:

agriculturenatural language processing

Type de record:

paperdatasetmodelsoftware
Créateur:
YanCheFenZhang, Yu
Hôte:avatar
Despite rapid advances in multimodal large language models, agricultural applications remain constrained by the scarcity of domain-tailored models, curated vision-language corpora, and rigorous evaluation. To address these challenges, we present the AgriGPT-VL Suite, a unified multimodal framework for agriculture. Our contributions are threefold. First, we introduce Agri-3M-VL, the largest vision-language corpus for agriculture to our knowledge, curated by a scalable multi-agent data generator; it comprises 1M image-caption pairs, 2M image-grounded VQA pairs, 50K expert-level VQA instances, and 15K GRPO reinforcement learning samples. Second, we develop AgriGPT-VL, an agriculture-specialized vision-language model trained via a progressive curriculum of textual grounding, multimodal shallow/deep alignment, and GRPO refinement. This method achieves strong multimodal reasoning while preserving text-only capability. Third, we establish AgriBench-VL-4K, a compact yet challenging evaluation suite with open-ended and image-grounded questions, paired with multi-metric evaluation and an LLM-as-a-judge framework. Experiments show that AgriGPT-VL outperforms leading general-purpose VLMs on AgriBench-VL-4K, achieving higher pairwise win rates in the LLM-as-a-judge evaluation. Meanwhile, it remains competitive on the text-only AgriBench-13K with no noticeable degradation of language ability. Ablation studies further confirm consistent gains from our alignment and GRPO refinement stages. We will open source all of the resources to support reproducible research and deployment in low-resource agricultural settings.

Visit

arxiv.org

Tags

Computation and Language

Similaires

AgriGPT-Omni: A Unified Speech-Vision-Text Framework for Multilingual Agricultural IntelligenceShaoza1/lesotho-vision-suiteBenchmarking Vision Language Models for Cultural UnderstandingJEEM: Vision-Language Understanding in Four Arabic DialectsEnhanced Arabic Document Understanding Using Vision-Language Modelselphassam5-bot/Towards-Improved-Understanding-of-VL-Distribution

AgriGPT-Omni: A Unified Speech-Vision-Text Framework for Multilingual Agricultural Intelligence

Despite rapid advances in multimodal large language models, agricultural applications remain constra

Shaoza1/lesotho-vision-suite

# 🏔️ MalotiTracker **Real-Time Color-Based Object Tracking for Lesotho Contexts** MalotiTracker is

Benchmarking Vision Language Models for Cultural Understanding

Foundation models and vision-language pre-training have notably advanced Vision Language Models (VLM

JEEM: Vision-Language Understanding in Four Arabic Dialects

We introduce JEEM, a benchmark designed to evaluate Vision-Language Models (VLMs) on visual understa

Enhanced Arabic Document Understanding Using Vision-Language Models

elphassam5-bot/Towards-Improved-Understanding-of-VL-Distribution

Towards Improved Understanding of Visceral Leishmaniasis Distribution in Kenya