paper-with-me

Papers

Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene Understanding

2024-09-05 · Yunze Man, Shuhong Zheng, Zhipeng Bao, Martial Hebert, Liang-Yan Gui, Yu-Xiong Wang

Complex 3D scene understanding has gained increasing attention, with scene encoding strategies playing a crucial role in this success. However, the optimal scene encoding strategies for various scenarios remain unclear, particularly compared to their image-based counterparts. To address this issue, we present a comprehensive study that probes various visual encoding models for 3D scene understanding, identifying the strengths and limitations of each model across different scenarios. Our evaluation spans seven vision foundation encoders, including image-based, video-based, and 3D foundation models. We evaluate these models in four tasks: Vision-Language Scene Reasoning, Visual Grounding, Segmentation, and Registration, each focusing on different aspects of scene understanding. Our evaluations yield key findings: DINOv2 demonstrates superior performance, video models excel in object-level tasks, diffusion models benefit geometric tasks, and language-pretrained models show unexpected limitations in language-related tasks. These insights challenge some conventional understandings, provide novel perspectives on leveraging visual foundation models, and highlight the need for more flexible encoder selection in future vision-language and scene-understanding tasks. Code: https://github.com/YunzeMan/Lexicon3D

📄 PDF Abstract BibTeX arXiv:2409.03757

Code (1)

yunzeman/lexicon3d pytorch

Tasks

Question AnsweringScene UnderstandingVisual Grounding

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Probing the 3D Awareness of Visual Foundation Models

2024-04-12 · CVPR 2024 1 · Mohamed El Banani, Amit Raj, Kevis-Kokitsi Maninis, Abhishek Kar 외

Recent advances in large-scale pretraining have yielded visual foundation models with strong capabilities. Not only can recent models generalize to arbitrary images for their training task, their intermediate representat…

What Do EEG Foundation Models Capture from Human Brain Signals?

2026-05-12 · Ling Tang, Qian Chen, Jilin Mei, Houshi Xu 외 arxiv

Clinical electroencephalogram (EEG) analysis rests on a hand-crafted feature catalog refined over decades, \emph{e.g.,} band power, connectivity, complexity, and more. Modern EEG foundation models bypass this catalog, le…

At the Lower End of Language---Exploring the Vulgar and Obscene Side of German

2019-08-01 · WS 2019 8 · Elisabeth Eder, Ulrike Krieg-Holz, Udo Hahn

In this paper, we describe a workflow for the data-driven acquisition and semantic scaling of a lexicon that covers lexical items from the lower end of the German language register{---}terms typically considered as rough…

A Lexicon for Profane and Obscene Text Identification in Bengali

2021-09-01 · RANLP 2021 9 · Salim Sazzed

Bengali is a low-resource language that lacks tools and resources for profane and obscene textual content detection. Until now, no lexicon exists for detecting obscenity in Bengali social media text. This study introduce…

POS

Improving 2D Feature Representations by 3D-Aware Fine-Tuning

2024-07-29 · Yuanwen Yue, Anurag Das, Francis Engelmann, Siyu Tang 외

Current visual foundation models are trained purely on unstructured 2D data, limiting their understanding of 3D structure of objects and scenes. In this work, we show that fine-tuning on 3D-aware data improves the qualit…

Depth EstimationSemantic Segmentation