paper-with-me

홈 › Papers

Refining Skewed Perceptions in Vision-Language Models through Visual Representations

2024-05-22 · Haocheng Dai, Sarang Joshi

Large vision-language models (VLMs), such as CLIP, have become foundational, demonstrating remarkable success across a variety of downstream tasks. Despite their advantages, these models, akin to other foundational systems, inherit biases from the disproportionate distribution of real-world data, leading to misconceptions about the actual environment. Prevalent datasets like ImageNet are often riddled with non-causal, spurious correlations that can diminish VLM performance in scenarios where these contextual elements are absent. This study presents an investigation into how a simple linear probe can effectively distill task-specific core features from CLIP's embedding for downstream applications. Our analysis reveals that the CLIP text representations are often tainted by spurious correlations, inherited in the biased pre-training dataset. Empirical evidence suggests that relying on visual representations from CLIP, as opposed to text embedding, is more practical to refine the skewed perceptions in VLMs, emphasizing the superior utility of visual representations in overcoming embedded biases. Our codes will be available here.

📄 PDF Abstract BibTeX arXiv:2405.14030

Code (0)

등록된 구현이 없습니다.

Tasks

Misconceptions

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Task Supportive and Personalized Human-Large Language Model Interaction: A User Study

2024-02-09 · Ben Wang, Jiqun Liu, Jamshed Karimnazarov, Nicolas Thompson

Large language model (LLM) applications, such as ChatGPT, are a powerful tool for online information-seeking (IS) and problem-solving tasks. However, users still face challenges initializing and refining prompts, and the…

Information RetrievalLanguage ModelingLanguage ModellingLarge Language Model+2

Social and Genetic Ties Drive Skewed Cross-Border Media Coverage of Disasters

2025-01-13 · Thiemo Fetzer, Prashant Garg

Climate change is increasing the frequency and severity of natural disasters worldwide. Media coverage of these events may be vital to generate empathy and mobilize global populations to address the common threat posed b…

Articles

Learning Through AI-Clones: Enhancing Self-Perception and Presentation Performance

2023-10-23 · Qingxiao Zheng, Zhuoer Chen, Yun Huang

This study examines the impact of AI-generated digital clones with self-images on enhancing perceptions and skills in online presentations. A mixed-design experiment with 44 international students compared self-recording…

Face SwappingVoice Cloning

How Susceptible are Large Language Models to Ideological Manipulation?

2024-02-18 · Kai Chen, Zihao He, Jun Yan, Taiwei Shi 외

Large Language Models (LLMs) possess the potential to exert substantial influence on public perceptions and interactions with information. This raises concerns about the societal impact that could arise if the ideologies…

Owls are wise and foxes are unfaithful: Uncovering animal stereotypes in vision-language models

2025-01-21 · Tabinda Aman, Mohammad Nadeem, Shahab Saquib Sohail, Mohammad Anas 외

Animal stereotypes are deeply embedded in human culture and language. They often shape our perceptions and expectations of various species. Our study investigates how animal stereotypes manifest in vision-language models…

Image Generation