Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Not Just What's There: Enabling CLIP to Comprehend Negated Visual Descriptions Without Fine-tuning

Record type:

papermodelsoftware
Creator:
XiaWu,LinChe
Host:avatar
Vision-Language Models (VLMs) like CLIP struggle to understand negation, often embedding affirmatives and negatives similarly (e.g., matching "no dog" with dog images). Existing methods refine negation understanding via fine-tuning CLIP's text encoder, risking overfitting. In this work, we propose CLIPGlasses, a plug-and-play framework that enhances CLIP's ability to comprehend negated visual descriptions. CLIPGlasses adopts a dual-stage design: a Lens module disentangles negated semantics from text embeddings, and a Frame module predicts context-aware repulsion strength, which is integrated into a modified similarity computation to penalize alignment with negated semantics, thereby reducing false positive matches. Experiments show that CLIP equipped with CLIPGlasses achieves competitive in-domain performance and outperforms state-of-the-art methods in cross-domain generalization. Its superiority is especially evident under low-resource conditions, indicating stronger robustness across domains.

Visit

arxiv.org

Tasks

image-text retrievalcomputer vision

Tags

Computer Vision and Pattern RecognitionMultimedia

Similar

Fine-Tuning Without Forgetting via Loss-Adaptive Learning RatesMultilingual Large Language Models do not comprehend all natural languages to equal degreesministercmanga/mBART50-fine-tuningMULTILINGUAL ADAPTIVE FINE-TUNING (MAFT)Abdelelta/Luganda-whisper-fine-tuningaman3013/Fine-tuning-Amharic-NER

Fine-Tuning Without Forgetting via Loss-Adaptive Learning Rates

Fine-tuning large language models on new data improves task performance but degrades capabilities le

Multilingual Large Language Models do not comprehend all natural languages to equal degrees

Large Language Models (LLMs) play a critical role in how humans access information. While their core

ministercmanga/mBART50-fine-tuning

Fine-tuning a pre-trained mBART50 Large Language Model for translating South African languages Proj

MULTILINGUAL ADAPTIVE FINE-TUNING (MAFT)

We introduce MAFT as an approach to adapt a multi-lingual PLM to a new set of languages. Adapting PLMs has been shown to be effective when adapting to a new domain (Gururangan et al., 2020) or language (Pfeiffer et al., 2020; Alabi et al., 2020; Adelani et al., 202

Abdelelta/Luganda-whisper-fine-tuning

aman3013/Fine-tuning-Amharic-NER

# Fine-tuning-Amharic-NER ## Overview EthioMart aims to become the primary hub for Telegram-based