paper-with-me

홈 › Papers

VGLD: Visually-Guided Linguistic Disambiguation for Monocular Depth Scale Recovery

2025-05-05 · Bojin Wu, Jing Chen

We propose a robust method for monocular depth scale recovery. Monocular depth estimation can be divided into two main directions: (1) relative depth estimation, which provides normalized or inverse depth without scale information, and (2) metric depth estimation, which involves recovering depth with absolute scale. To obtain absolute scale information for practical downstream tasks, utilizing textual information to recover the scale of a relative depth map is a highly promising approach. However, since a single image can have multiple descriptions from different perspectives or with varying styles, it has been shown that different textual descriptions can significantly affect the scale recovery process. To address this issue, our method, VGLD, stabilizes the influence of textual information by incorporating high-level semantic information from the corresponding image alongside the textual description. This approach resolves textual ambiguities and robustly outputs a set of linear transformation parameters (scalars) that can be globally applied to the relative depth map, ultimately generating depth predictions with metric-scale accuracy. We validate our method across several popular relative depth models(MiDas, DepthAnything), using both indoor scenes (NYUv2) and outdoor scenes (KITTI). Our results demonstrate that VGLD functions as a universal alignment module when trained on multiple datasets, achieving strong performance even in zero-shot scenarios. Code is available at: https://github.com/pakinwu/VGLD.

📄 PDF Abstract BibTeX arXiv:2505.02704

Code (1)

pakinwu/vgld 공식 구현

Tasks

Depth EstimationMonocular Depth Estimation

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

VIDA: A dataset for Visually Dependent Ambiguity in Multimodal Machine Translation

2026-05-03 · Jingheng Pan, Xintong Wang, Longyue Wang, Liang Ding 외 arxiv

Ambiguity resolution is a key challenge in multimodal machine translation (MMT), where models must genuinely leverage visual input to map an ambiguous expression to its intended meaning. Although prior work has proposed …

Multimodal Machine Translation

From Speech-to-Spatial: Grounding Utterances on A Live Shared View with Augmented Reality

2026-02-03 · Yoonsang Kim, Divyansh Pradhan, Devshree Jadeja, Arie Kaufman arxiv

We introduce Speech-to-Spatial, a referent disambiguation framework that converts verbal remote-assistance instructions into spatially grounded AR guidance. Unlike prior systems that rely on additional cues (e.g., gestur…

Psycholinguistics, Lexicography, and Word Sense Disambiguation

2012-11-01 · PACLIC 2012 11 · Oi Yee Kwong
Word Sense Disambiguation

Part-of-Speech Tag Disambiguation by Cross-Linguistic Majority Vote

2014-08-01 · WS 2014 8 · No{\"e}mi Aepli, Ruprecht von Waldenfels, Tanja Samard{\v{z}}i{\'c}
Machine TranslationPart-Of-Speech TaggingTAGWord Alignment

ID-VTG: Image-Disambiguated Video Temporal Grounding

2026-08-20 · Minghang Zheng, Jingli Wei, Hongyi Yang, Yang Liu arxiv

Video Temporal Grounding (VTG) faces significant challenges when natural language queries must distinguish between multiple events involving visually similar entities, particularly when relying on fine-grained visual att…

Natural Language Queries