paper-with-me

Papers

Complete 3d relationships extraction modality alignment network for 3d dense captioning

2024-08-01 · IEEE Transactions on Visualization and Computer Graphics. 2024 8 · Aihua Mao, Zhi Yang, Wanxin Chen, Ran Yi, Yong-Jin Liu

3D dense captioning aims to semantically describe each object detected in a 3D scene, which plays a significant role in 3D scene understanding. Previous works lack a complete definition of 3D spatial relationships and the directly integrate visual and language modalities, thus ignoring the discrepancies between the two modalities. To address these issues, we propose a novel complete 3D relationship extraction modality alignment network, which consists of three steps: 3D object detection, complete 3D relationships extraction, and modality alignment caption. To comprehensively capture the 3D spatial relationship features, we define a complete set of 3D spatial relationships, including the local spatial relationship between objects and the global spatial relationship between each object and the entire scene. To this end, we propose a complete 3D relationships extraction module based on message passing and self-attention to mine multi-scale spatial relationship features and inspect the transformation to obtain features in different views. In addition, we propose the modality alignment caption module to fuse multi-scale relationship features and generate descriptions to bridge the semantic gap from the visual space to the language space with the prior information in the word embedding, and help generate improved descriptions for the 3D scene. Extensive experiments demonstrate that the proposed model outperforms the state-of-the-art methods on the ScanRefer and Nr3D datasets.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

3D dense captioning3D Object DetectionDense Captioningobject-detectionObject DetectionScene Understanding

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Visually Guided Spatial Relation Extraction from Text

2018-06-01 · NAACL 2018 6 · Taher Rahgooy, Umar Manzoor, Parisa Kordjamshidi

Extraction of spatial relations from sentences with complex/nesting relationships is very challenging as often needs resolving inherent semantic ambiguities. We seek help from visual modality to fill the information gap …

Activity RecognitionImage CaptioningImage RetrievalObject Localization+4

Dense Multimodal Alignment for Open-Vocabulary 3D Scene Understanding

2024-07-13 · Ruihuang Li, Zhengqiang Zhang, Chenhang He, Zhiyuan Ma 외

Recent vision-language pre-training models have exhibited remarkable generalization ability in zero-shot recognition tasks. Previous open-vocabulary 3D scene understanding methods mostly focus on training 3D models using…

Scene UnderstandingZero-Shot Learning

Fixed-Length Dense Fingerprint Representation

2025-05-06 · Zhiyu Pan, Xiongjun Guan, Yongjie Duan, Jianjiang Feng 외

Fixed-length fingerprint representations, which map each fingerprint to a compact and fixed-size feature vector, are computationally efficient and well-suited for large-scale matching. However, designing a robust represe…

CrossMAE: Cross-Modality Masked Autoencoders for Region-Aware Audio-Visual Pre-Training

2024-01-01 · CVPR 2024 1 · Yuxin Guo, Siyang Sun, Shuailei Ma, Kecheng Zheng 외

Learning joint and coordinated features across modalities is essential for many audio-visual tasks. Existing pre-training methods primarily focus on global information neglecting fine-grained features and positions l…

Robust Incomplete-Modality Alignment for Ophthalmic Disease Grading and Diagnosis via Labeled Optimal Transport

2025-07-07 · Qinkai Yu, Jianyang Xie, Yitian Zhao, Cheng Chen 외 arxiv

Multimodal ophthalmic imaging-based diagnosis integrates color fundus image with optical coherence tomography (OCT) to provide a comprehensive view of ocular pathologies. However, the uneven global distribution of health…