paper-with-me

홈 › Papers

Pixel-Aligned Language Model

2024-01-01 · CVPR 2024 1 · Jiarui Xu, Xingyi Zhou, Shen Yan, Xiuye Gu, Anurag Arnab, Chen Sun, Xiaolong Wang, Cordelia Schmid

Large language models have achieved great success in recent years so as their variants in vision. Existing vision-language models can describe images in natural languages answer visual-related questions or perform complex reasoning about the image. However it is yet unclear how localization tasks such as word grounding or referring localization can be performed using large language models. In this work we aim to develop a vision-language model that can take locations for example a set of points or boxes as either inputs or outputs. When taking locations as inputs the model performs location-conditioned captioning which generates captions for the indicated object or region. When generating locations as outputs our model regresses pixel coordinates for each output word generated by the language model and thus performs dense word grounding. Our model is pre-trained on the Localized Narrative dataset which contains pixel-word-aligned captioning from human attention. We show our model can be applied to various location-aware vision-language tasks including referring localization location-conditioned captioning and dense object captioning archiving state-of-the-art performance on RefCOCO and Visual Genome.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modellingmodel

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

PLAF: Pixel-wise Language-Aligned Feature Extraction for Efficient 3D Scene Understanding

2026-04-17 · Junjie Wen, Junlin He, Fei Ma, Jinqiang Cui arxiv

Accurate open-vocabulary 3D scene understanding requires semantic representations that are both language-aligned and spatially precise at the pixel level, while remaining scalable when lifted to 3D space. However, existi…

Scene Understanding

Pixel Aligned Language Models

2023-12-14 · Jiarui Xu, Xingyi Zhou, Shen Yan, Xiuye Gu 외

Large language models have achieved great success in recent years, so as their variants in vision. Existing vision-language models can describe images in natural languages, answer visual-related questions, or perform com…

Language ModelingLanguage Modelling

NOVA3R: Non-pixel-aligned Visual Transformer for Amodal 3D Reconstruction

2026-03-04 · Weirong Chen, Chuanxia Zheng, Ganlin Zhang, Andrea Vedaldi 외 arxiv

We present NOVA3R, an effective approach for non-pixel-aligned 3D reconstruction from a set of unposed images in a feed-forward manner. Unlike pixel-aligned methods that tie geometry to per-ray predictions, our formulati…

3D ReconstructionPoint Clouds

Pixel-aligned RGB-NIR Stereo Imaging and Dataset for Robot Vision

2024-11-27 · CVPR 2025 1 · Jinnyeong Kim, Seung-Hwan Baek

Integrating RGB and NIR stereo imaging provides complementary spectral information, potentially enhancing robotic 3D vision in challenging lighting conditions. However, existing datasets and imaging systems lack pixel-le…

PVSeRF: Joint Pixel-, Voxel- and Surface-Aligned Radiance Field for Single-Image Novel View Synthesis

2022-02-10 · Xianggang Yu, Jiapeng Tang, Yipeng Qin, Chenghong Li 외

We present PVSeRF, a learning framework that reconstructs neural radiance fields from single-view RGB images, for novel view synthesis. Previous solutions, such as pixelNeRF, rely only on pixel-aligned features and suffe…

DisentanglementNovel View Synthesis