paper-with-me

홈 › Papers

Vision-Language Integration for Zero-Shot Scene Understanding in Real-World Environments

2025-10-29 · Manjunath Prasad Holenarasipura Rajiv, B. M. Vidyavathi arxiv

Zero-shot scene understanding in real-world settings presents major challenges due to the complexity and variability of natural scenes, where models must recognize new objects, actions, and contexts without prior labeled examples. This work proposes a vision-language integration framework that unifies pre-trained visual encoders (e.g., CLIP, ViT) and large language models (e.g., GPT-based architectures) to achieve semantic alignment between visual and textual modalities. The goal is to enable robust zero-shot comprehension of scenes by leveraging natural language as a bridge to generalize over unseen categories and contexts. Our approach develops a unified model that embeds visual inputs and textual prompts into a shared space, followed by multimodal fusion and reasoning layers for contextual interpretation. Experiments on Visual Genome, COCO, ADE20K, and custom real-world datasets demonstrate significant gains over state-of-the-art zero-shot models in object recognition, activity detection, and scene captioning. The proposed system achieves up to 18% improvement in top-1 accuracy and notable gains in semantic coherence metrics, highlighting the effectiveness of cross-modal alignment and language grounding in enhancing generalization for real-world scene understanding.

📄 PDF Abstract BibTeX arXiv:2510.25070

Code (0)

등록된 구현이 없습니다.

Tasks

Scene UnderstandingActivity DetectionObject Recognition

Similar Papers 제목 키워드 기반

VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding

2024-10-17 · Runsen Xu, Zhiwei Huang, Tai Wang, Yilun Chen 외

3D visual grounding is crucial for robots, requiring integration of natural language and 3D scene understanding. Traditional methods depending on supervised learning with 3D point clouds are limited by scarce datasets. R…

3D geometry3D visual groundingObjectScene Understanding+1

Dynamic Context-Aware Scene Reasoning Using Vision-Language Alignment in Zero-Shot Real-World Scenarios

2025-10-30 · Manjunath Prasad Holenarasipura Rajiv, B. M. Vidyavathi arxiv

In real-world environments, AI systems often face unfamiliar scenarios without labeled data, creating a major challenge for conventional scene understanding models. The inability to generalize across unseen contexts limi…

Zero-shot GeneralizationScene Understanding

Enhancing Remote Sensing Vision-Language Models for Zero-Shot Scene Classification

2024-09-01 · Karim El Khoury, Maxime Zanella, Benoît Gérin, Tiffanie Godelaine 외

Vision-Language Models for remote sensing have shown promising uses thanks to their extensive pretraining. However, their conventional usage in zero-shot scene classification methods still involves dividing large images …

Scene ClassificationTransductive Zero-Shot Classificationzero-shot-classificationZero-Shot Learning

Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN

2026-05-14 · Ziyi Xia, Chaoran Xiong, Litao Wei, Xinhao Hu 외 arxiv

Zero-shot vision-and-language navigation (VLN) has gained significant attention due to its minimal data collection costs and inherent generalization. This paradigm is typically driven by the integration of pre-trained Vi…

Scene Understanding

Leveraging Large (Visual) Language Models for Robot 3D Scene Understanding

2022-09-12 · William Chen, Siyi Hu, Rajat Talak, Luca Carlone

Abstract semantic 3D scene understanding is a problem of critical importance in robotics. As robots still lack the common-sense knowledge about household objects and locations of an average human, we investigate the use …

Common Sense ReasoningScene ClassificationScene Understanding