paper-with-me

홈 › Papers

Exploring the Zero-Shot Capabilities of Vision-Language Models for Improving Gaze Following

2024-06-06 · Anshul Gupta, Pierre Vuillecard, Arya Farkhondeh, Jean-Marc Odobez

Contextual cues related to a person's pose and interactions with objects and other people in the scene can provide valuable information for gaze following. While existing methods have focused on dedicated cue extraction methods, in this work we investigate the zero-shot capabilities of Vision-Language Models (VLMs) for extracting a wide array of contextual cues to improve gaze following performance. We first evaluate various VLMs, prompting strategies, and in-context learning (ICL) techniques for zero-shot cue recognition performance. We then use these insights to extract contextual cues for gaze following, and investigate their impact when incorporated into a state of the art model for the task. Our analysis indicates that BLIP-2 is the overall top performing VLM and that ICL can improve performance. We also observe that VLMs are sensitive to the choice of the text prompt although ensembling over multiple text prompts can provide more robust performance. Additionally, we discover that using the entire image along with an ellipse drawn around the target person is the most effective strategy for visual prompting. For gaze following, incorporating the extracted cues results in better generalization performance, especially when considering a larger set of cues, highlighting the potential of this approach.

📄 PDF Abstract BibTeX arXiv:2406.03907

Code (0)

등록된 구현이 없습니다.

Tasks

In-Context LearningVisual Prompting

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Good Questions Help Zero-Shot Image Reasoning

2023-12-04 · Kaiwen Yang, Tao Shen, Xinmei Tian, Xiubo Geng 외

Aligning the recent large language models (LLMs) with computer vision models leads to large vision-language models (LVLMs), which have paved the way for zero-shot image reasoning tasks. However, LVLMs are usually trained…

Fine-Grained Image ClassificationQuestion AnsweringVisual EntailmentVisual Question Answering

Exploring Vision-Language Models for Open-Vocabulary Zero-Shot Action Segmentation

2026-02-24 · Asim Unmesh, Kaki Ramesh, Mayank Patel, Rahul Jain 외 arxiv

Temporal Action Segmentation (TAS) requires dividing videos into action segments, yet the vast space of activities and alternative breakdowns makes collecting comprehensive datasets infeasible. Existing methods remain li…

Action Segmentation

TINA: Think, Interaction, and Action Framework for Zero-Shot Vision Language Navigation

2024-03-13 · Dingbang Li, Wenzhou Chen, Xin Lin

Zero-shot navigation is a critical challenge in Vision-Language Navigation (VLN) tasks, where the ability to adapt to unfamiliar instructions and to act in unknown environments is essential. Existing supervised learning-…

Question AnsweringVision-Language Navigation

VL-Taboo: An Analysis of Attribute-based Zero-shot Capabilities of Vision-Language Models

2022-09-12 · Felix Vogel, Nina Shvetsova, Leonid Karlinsky, Hilde Kuehne

Vision-language models trained on large, randomly collected data had significant impact in many areas since they appeared. But as they show great performance in various fields, such as image-text-retrieval, their inner w…

AttributeImage-text RetrievalRetrievalText Retrieval+1

Vision-Language Models Performing Zero-Shot Tasks Exhibit Gender-based Disparities

2023-01-26 · Melissa Hall, Laura Gustafson, Aaron Adcock, Ishan Misra 외

We explore the extent to which zero-shot vision-language models exhibit gender bias for different vision tasks. Vision models traditionally required task-specific labels for representing concepts, as well as finetuning; …

image-classificationImage Classificationobject-detectionObject Detection+3