paper-with-me

Papers

Composition Vision-Language Understanding via Segment and Depth Anything Model

2024-06-07 · Mingxiao Huo, Pengliang Ji, Haotian Lin, Junchen Liu, Yixiao Wang, Yijun Chen

We introduce a pioneering unified library that leverages depth anything, segment anything models to augment neural comprehension in language-vision model zero-shot understanding. This library synergizes the capabilities of the Depth Anything Model (DAM), Segment Anything Model (SAM), and GPT-4V, enhancing multimodal tasks such as vision-question-answering (VQA) and composition reasoning. Through the fusion of segmentation and depth analysis at the symbolic instance level, our library provides nuanced inputs for language models, significantly advancing image interpretation. Validated across a spectrum of in-the-wild real-world images, our findings showcase progress in vision-language models through neural-symbolic integration. This novel approach melds visual and language analysis in an unprecedented manner. Overall, our library opens new directions for future research aimed at decoding the complexities of the real world through advanced multimodal technologies and our code is available at \url{https://github.com/AnthonyHuo/SAM-DAM-for-Compositional-Reasoning}.

📄 PDF Abstract BibTeX arXiv:2406.18591

Code (1)

anthonyhuo/sam-dam-for-compositional-reasoning 공식 구현 pytorch

Tasks

Question AnsweringVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

Library 설명 없음

Similar Papers 제목 키워드 기반

Perceptio: Perception Enhanced Vision Language Models via Spatial Token Generation

2026-03-19 · Yuchen Li, Amanmeet Garg, Shalini Chaudhuri, Rui Zhao 외 arxiv

Large Vision Language Models (LVLMs) excel at semantic understanding but struggle with fine grained spatial grounding, as the model must implicitly infer complex geometry without ever producing a spatial interpretation. …

Referring Expression SegmentationSemantic SegmentationSpatial Reasoning

From Recovery to Drop-off: How Action Post-training Reduces a VLM's Late-Layer Depth Decodability

2026-08-09 · Alexander Hackett, Arnaud Denis-Remillard, Axel Cassou arxiv

How much of a vision-language model's (VLM) spatial understanding remains after the action post-training process of building a vision-language-action model (VLA)? We probe depth perception, a primitive of spatiogeometric…

Depth-wise layering of 3d images using dense depth maps: a threshold based approach

2020-10-05 · Seyedsaeid Mirkamali, P. Nagabhushan

Image segmentation has long been a basic problem in computer vision. Depth-wise Layering is a kind of segmentation that slices an image in a depth-wise sequence unlike the conventional image segmentation problems dealing…

Image SegmentationSegmentationSemantic Segmentation

Image Generators are Generalist Vision Learners

2026-04-22 · Valentin Gabeur, Shangbang Long, Songyou Peng, Paul Voigtlaender 외 arxiv

Recent works show that image and video generators exhibit zero-shot visual understanding behaviors, in a way reminiscent of how LLMs develop emergent capabilities of language understanding and reasoning from generative p…

Depth EstimationImage GenerationText Generation

Uni3DL: Unified Model for 3D and Language Understanding

2023-12-05 · Xiang Li, Jian Ding, Zhaoyang Chen, Mohamed Elhoseiny

In this work, we present Uni3DL, a unified model for 3D and Language understanding. Distinct from existing unified vision-language models in 3D which are limited in task variety and predominantly dependent on projected m…

Cross-Modal RetrievalInstance Segmentationmodelobject-detection+4