paper-with-me

홈 › Papers

DOMR: Establishing Cross-View Segmentation via Dense Object Matching

2025-08-06 · Jitong Liao, Yulu Gao, Shaofei Huang, Jialin Gao, Jie Lei, Ronghua Liang, Si Liu arxiv

Cross-view object correspondence involves matching objects between egocentric (first-person) and exocentric (third-person) views. It is a critical yet challenging task for visual understanding. In this work, we propose the Dense Object Matching and Refinement (DOMR) framework to establish dense object correspondences across views. The framework centers around the Dense Object Matcher (DOM) module, which jointly models multiple objects. Unlike methods that directly match individual object masks to image features, DOM leverages both positional and semantic relationships among objects to find correspondences. DOM integrates a proposal generation module with a dense matching module that jointly encodes visual, spatial, and semantic cues, explicitly constructing inter-object relationships to achieve dense matching among objects. Furthermore, we combine DOM with a mask refinement head designed to improve the completeness and accuracy of the predicted masks, forming the complete DOMR framework. Extensive evaluations on the Ego-Exo4D benchmark demonstrate that our approach achieves state-of-the-art performance with a mean IoU of 49.7% on Ego$\to$Exo and 55.2% on Exo$\to$Ego. These results outperform those of previous methods by 5.8% and 4.3%, respectively, validating the effectiveness of our integrated approach for cross-view understanding.

📄 PDF Abstract BibTeX arXiv:2508.04050

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

VisDoM: Multi-Document QA with Visually Rich Elements Using Multimodal Retrieval-Augmented Generation

2024-12-14 · Manan Suri, Puneet Mathur, Franck Dernoncourt, Kanika Goswami 외

Understanding information from a collection of multiple documents, particularly those with visually rich elements, is important for document-grounded question answering. This paper introduces VisDoMBench, the first compr…

Question AnsweringRAGRetrievalRetrieval-augmented Generation

Towards Establishing Dense Correspondence on Multiview Coronary Angiography: From Point-to-Point to Curve-to-Curve Query Matching

2023-12-18 · Yifan Wu, Rohit Jena, Mehmet Gulsun, Vivek Singh 외

Coronary angiography is the gold standard imaging technique for studying and diagnosing coronary artery disease. However, the resulting 2D X-ray projections lose 3D information and exhibit visual ambiguities. In this wor…

MV-RoMa: From Pairwise Matching into Multi-View Track Reconstruction

2026-03-29 · Jongmin Lee, Seungyeop Kang, Sungjoo Yoo arxiv

Establishing consistent correspondences across images is essential for 3D vision tasks such as structure-from-motion (SfM), yet most existing matchers operate in a pairwise manner, often producing fragmented and geometri…

Training-free Detection of AI-generated images via Cropping Robustness

2025-11-18 · Sungik Choi, Hankook Lee, Moontae Lee arxiv

AI-generated image detection has become crucial with the rapid advancement of vision-generative models. Instead of training detectors tailored to specific datasets, we study a training-free approach leveraging self-super…

DCD: A Semantic Segmentation Model for Fetal Ultrasound Four-Chamber View

2025-06-10 · Donglian Li, Hui Guo, Minglang Chen, Huizhen Chen 외

Accurate segmentation of anatomical structures in the apical four-chamber (A4C) view of fetal echocardiography is essential for early diagnosis and prenatal evaluation of congenital heart disease (CHD). However, precise …

SegmentationSemantic Segmentation