Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Incorporating Geo-Diverse Knowledge into Prompting for Increased Geographical Robustness in Object Recognition

Domain:

natural language processing

Record type:

paper
Creator:
BueMalLi,Kov
Host:avatar
Existing object recognition models have been shown to lack robustness in diverse geographical scenarios due to domain shifts in design and context. Class representations need to be adapted to more accurately reflect an object concept under these shifts. In the absence of training data from target geographies, we hypothesize that geographically diverse descriptive knowledge of categories can enhance robustness. For this purpose, we explore the feasibility of probing a large language model for geography-based object knowledge, and we examine the effects of integrating knowledge into zero-shot and learnable soft prompting with CLIP. Within this exploration, we propose geography knowledge regularization to ensure that soft prompts trained on a source set of geographies generalize to an unseen target set. Accuracy gains over prompting baselines on DollarStreet while training only on Europe data are up to +2.8/1.2/1.6 on target data from Africa/Asia/Americas, and +4.6 overall on the hardest classes. Competitive performance is shown vs. few-shot target training, and analysis is provided to direct future study of geographical robustness. To appear in IEEE/CVF Computer Vision and Pattern Recognition Conference (CVPR), 2024

Visit

arxiv.org

Tags

Computer Vision and Pattern RecognitionArtificial IntelligenceMachine Learning

Similar

Rapid Language Adaptation for Multilingual E2E Speech Recognition Using Encoder PromptingBroaden the Vision: Geo-Diverse Visual Commonsense ReasoningIncorporating territory compression into population modelsObject Recognition for Economic Development from Daytime Satellite ImageryIncorporating evolutionary history into conservation planning in biodiversity hotspotsIncorporating Telementorship Into Laboratory Capacity Building Initiatives for Improved AMR Surveillance in Ethiopia

Rapid Language Adaptation for Multilingual E2E Speech Recognition Using Encoder Prompting

End-to-end multilingual speech recognition models handle multiple languages through a single model,

Broaden the Vision: Geo-Diverse Visual Commonsense Reasoning

Commonsense is defined as the knowledge on which everyone agrees. However, certain types of commonsense knowledge are correlated with culture and geographic locations and they are only shared locally. For example, the scenes of wedding ceremonies vary across region

Incorporating territory compression into population models

The ideal despotic distribution, whereby the lifetime reproductive success a territory's owner achie

Object Recognition for Economic Development from Daytime Satellite Imagery

Reliable data about the stock of physical capital and infrastructure in developing countries is typi

Incorporating evolutionary history into conservation planning in biodiversity hotspots

There is increased evidence that incorporating evolutionary history directly in conservation actions

Incorporating Telementorship Into Laboratory Capacity Building Initiatives for Improved AMR Surveillance in Ethiopia

Background: In July 2017, recognizing the threat that antimicrobial resistance poses to the populat