paper-with-me

홈 › Papers

Bridging Geometric and Semantic Foundation Models for Generalized Monocular Depth Estimation

2025-05-29 · Sanggyun Ma, Wonjoon Choi, Jihun Park, Jaeyeul Kim, Seunghun Lee, Jiwan Seo, Sunghoon Im

We present Bridging Geometric and Semantic (BriGeS), an effective method that fuses geometric and semantic information within foundation models to enhance Monocular Depth Estimation (MDE). Central to BriGeS is the Bridging Gate, which integrates the complementary strengths of depth and segmentation foundation models. This integration is further refined by our Attention Temperature Scaling technique. It finely adjusts the focus of the attention mechanisms to prevent over-concentration on specific features, thus ensuring balanced performance across diverse inputs. BriGeS capitalizes on pre-trained foundation models and adopts a strategy that focuses on training only the Bridging Gate. This method significantly reduces resource demands and training time while maintaining the model's ability to generalize effectively. Extensive experiments across multiple challenging datasets demonstrate that BriGeS outperforms state-of-the-art methods in MDE for complex scenes, effectively handling intricate structures and overlapping objects.

📄 PDF Abstract BibTeX arXiv:2505.23400

Code (0)

등록된 구현이 없습니다.

Tasks

Depth EstimationMonocular Depth Estimation

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Focus 설명 없음

Similar Papers 제목 키워드 기반

Feature4X: Bridging Any Monocular Video to 4D Agentic AI with Versatile Gaussian Feature Fields

2025-03-26 · CVPR 2025 1 · Shijie Zhou, Hui Ren, Yijia Weng, Shuwang Zhang 외

Recent advancements in 2D and multimodal models have achieved remarkable success by leveraging large-scale training on extensive datasets. However, extending these achievements to enable free-form interactions and high-l…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Pseudo-Generalized Dynamic View Synthesis from a Video

2023-10-12 · Xiaoming Zhao, Alex Colburn, Fangchang Ma, Miguel Angel Bautista 외

Rendering scenes observed in a monocular video from novel viewpoints is a challenging problem. For static scenes the community has studied both scene-specific optimization techniques, which optimize on every test scene, …

Novel View Synthesis

Geometry-Consistent Endoscopic Representations for Image-Guided Navigation via Structured Foundation Model Adaptation

2026-06-15 · Hongchao Shu, Roger D. Soberanis-Mukul, Hao Ding, Morgan Ringel 외 arxiv

Accurate vision-based navigation in monocular endoscopy is difficult due to limited depth cues, weak tissue texture, non-rigid deformation, and substantial appearance variation across domains, all of which complicate pos…

Monocular Depth EstimationRepresentation LearningPose Estimation

MonoCLUE : Object-Aware Clustering Enhances Monocular 3D Object Detection

2025-11-11 · Sunghun Yang, Minhyeok Lee, Jungho Lee, Sangyoun Lee arxiv

Monocular 3D object detection offers a cost-effective solution for autonomous driving but suffers from ill-posed depth and limited field of view. These constraints cause a lack of geometric cues and reduced accuracy in o…

Monocular 3D Object DetectionAutonomous Driving

FMGS-Avatar: Mesh-Guided 2D Gaussian Splatting with Foundation Model Priors for 3D Monocular Avatar Reconstruction

2025-09-18 · Jinlong Fan, Bingyu Hu, Xingguang Li, Yuxiang Yang 외 arxiv

Reconstructing high-fidelity animatable human avatars from monocular videos remains challenging due to insufficient geometric information in single-view observations. While recent 3D Gaussian Splatting methods have shown…