paper-with-me

Papers

Explicitly Guided Information Interaction Network for Cross-modal Point Cloud Completion

2024-07-03 · Hang Xu, Chen Long, Wenxiao Zhang, YuAn Liu, Zhen Cao, Zhen Dong, Bisheng Yang

In this paper, we explore a novel framework, EGIInet (Explicitly Guided Information Interaction Network), a model for View-guided Point cloud Completion (ViPC) task, which aims to restore a complete point cloud from a partial one with a single view image. In comparison with previous methods that relied on the global semantics of input images, EGIInet efficiently combines the information from two modalities by leveraging the geometric nature of the completion task. Specifically, we propose an explicitly guided information interaction strategy supported by modal alignment for point cloud completion. First, in contrast to previous methods which simply use 2D and 3D backbones to encode features respectively, we unified the encoding process to promote modal alignment. Second, we propose a novel explicitly guided information interaction strategy that could help the network identify critical information within images, thus achieving better guidance for completion. Extensive experiments demonstrate the effectiveness of our framework, and we achieved a new state-of-the-art (+16% CD over XMFnet) in benchmark datasets despite using fewer parameters than the previous methods. The pre-trained model and code and are available at https://github.com/WHU-USI3DV/EGIInet.

📄 PDF Abstract BibTeX arXiv:2407.02887

Code (1)

whu-usi3dv/egiinet 공식 구현 pytorch

Tasks

Point Cloud Completion

Similar Papers 제목 키워드 기반

CIGTSurv: Clinical Information Guided Tri-modal Survival Prediction with Local Prototype Association and Global Feature Alignment

2026-08-04 · Jing Dai, Qibin Zhang, Weiwei Zhou, Mingde Xu 외 arxiv

Multimodal learning has significantly advanced survival prediction by integrating pathology images with genomic data. However, clinical information, despite its critical role in reflecting a patient' s overall health, re…

Cross-Modal Conditioned Reconstruction for Language-guided Medical Image Segmentation

2024-04-03 · Xiaoshuang Huang, Hongxiang Li, Meng Cao, Long Chen 외

Recent developments underscore the potential of textual information in enhancing learning models for a deeper understanding of medical visual semantics. However, language-guided medical image segmentation still faces a c…

Image SegmentationMedical Image SegmentationSegmentationSemantic Segmentation

Rethinking Modality Reliability in Multimodal Sentiment Analysis with Incomplete Observations

2026-08-04 · Chunlei Meng, Jacqueline J. Pang, Pengbin Feng, Zhenyu Yu 외 arxiv

Multimodal Sentiment Analysis (MSA) integrates text, audio, and vision to infer human affect, yet real-world multimodal observations are often incomplete. Existing methods for incomplete-observation MSA mainly follow two…

Multimodal Sentiment AnalysisRepresentation Learning

MI-Pruner: Crossmodal Mutual Information-guided Token Pruner for Efficient MLLMs

2026-04-03 · Jiameng Li, Aleksei Tiulpin, Matthew B. Blaschko arxiv

For multimodal large language models (MLLMs), visual information is relatively sparse compared with text. As a result, research on visual pruning emerges for efficient inference. Current approaches typically measure toke…

Orthogonalized Multimodal Contrastive Learning with Asymmetric Masking for Structured Representations

2026-02-16 · Carolin Cissee, Raneen Younis, Zahra Ahmadi arxiv

Multimodal learning seeks to integrate information from heterogeneous sources, where signals may be shared across modalities, specific to individual modalities, or emerge only through their interaction. While self-superv…

Contrastive Learning