paper-with-me

홈 › Papers

From Pixels to Prompts: Vision-Language Models

2026-05-08 · Khang Hoang Nhat Vo arxiv

When you read a paper about a new Vision-Language Model today, it can be easy to forget how strange this idea would have sounded not so long ago. Teaching machines to see was already hard. Teaching them to read and generate language was already hard. Asking them to do both at once - and then to reason, answer questions, follow instructions, and sometimes even surprise us - still carries a quiet trace of science fiction, even as it becomes routine. This book was born from a simple feeling: it is too easy to get lost. The field moves quickly, new model names appear constantly, and the gap between "I know the buzzwords" and "I actually understand how this works" can feel uncomfortably wide. I have felt that gap many times. If you are holding this book, you probably have too. My goal is not to provide an exhaustive catalog of every dataset, benchmark, and new model variant. Instead, I want to offer something more modest - and, I hope, more durable: a clear mental map of Vision-Language Models. Enough structure that you can read new papers with confidence; enough intuition that you can design your own systems without feeling as if you are assembling LEGO bricks blindly.

📄 PDF Abstract BibTeX arXiv:2605.07544

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Attention, Please! PixelSHAP Reveals What Vision-Language Models Actually Focus On

2025-03-09 · Roni Goldshmidt

Interpretability in Vision-Language Models (VLMs) is crucial for trust, debugging, and decision-making in high-stakes applications. We introduce PixelSHAP, a model-agnostic framework extending Shapley-based analysis to s…

Autonomous DrivingDecision Making

CARIS: Context-Augmented Referring Image Segmentation

2023-10-27 · ACM MM 2023 10 · Sun-Ao Liu, Yiheng Zhang, Zhaofan Qiu, Hongtao Xie 외

Referring image segmentation aims to segment the target object described by a natural-language utterance. Recent approaches typically distinguish pixels by aligning pixel-wise visual features with linguistic features ext…

DecoderImage SegmentationSegmentationSemantic Segmentation

PointSeg: A Training-Free Paradigm for 3D Scene Segmentation via Foundation Models

2024-03-11 · Qingdong He, Jinlong Peng, Zhengkai Jiang, Xiaobin Hu 외

Recent success of vision foundation models have shown promising performance for the 2D perception tasks. However, it is difficult to train a 3D foundation network directly due to the limited dataset and it remains under …

Scene Segmentation

Sequential Modeling Enables Scalable Learning for Large Vision Models

2023-12-01 · CVPR 2024 1 · Yutong Bai, Xinyang Geng, Karttikeya Mangalam, Amir Bar 외

We introduce a novel sequential modeling approach which enables learning a Large Vision Model (LVM) without making use of any linguistic data. To do this, we define a common format, "visual sentences", in which we can re…

Diversity

EV-CLIP: Efficient Visual Prompt Adaptation for CLIP in Few-shot Action Recognition under Visual Challenges

2026-04-24 · Hyo Jin Jon, Longbin Jin, Eun Yi Kim arxiv

CLIP has demonstrated strong generalization in visual domains through natural language supervision, even for video action recognition. However, most existing approaches that adapt CLIP for action recognition have primari…

Action Recognition