paper-with-me

홈 › Papers

GOMA: Toward Structure-Driven Multimodal Alignment from a Graph Signal Smoothing Perspective

2026-05-15 · Xu Wang, Xunkai Li, Yinlin Zhu, Rong-Hua Li, Guoren Wang arxiv

Multimodal alignment is commonly learned from isolated image-text pairs via CLIP-style dual encoders, leaving the relational context among entities largely unused. Multimodal attributed graphs (MAGs), where nodes carry multimodal attributes and edges encode corpus structure, provide a natural setting for refining frozen vision-language embeddings. This refinement is challenging: visual, textual, and cross-modal relations often induce different neighborhood geometries, while unrestricted graph propagation can quickly over-smooth retrieval representations. Effectively leveraging graph context therefore requires simultaneously breaking modality-specific topological barriers, controlling the smoothing regime, and preserving informative smoothing before semantic boundaries collapse. We propose Graph-Optimized Multimodal Alignment (GOMA), a structure-driven post-alignment framework that views frozen multimodal embeddings as graph signals and addresses these requirements through a unified retrieval-oriented design. GOMA decouples three key design choices: where messages should flow, how multimodal evidence should propagate, and which smoothing depth should be retained. Concretely, it learns modality-aware propagation operators, performs finite-step coupled smoothing without diagonal cross-modal shortcuts, and adaptively reads out node-specific smoothing trajectories to preserve useful smoothing before collapse. All experiments follow a transductive MAG retrieval protocol where the graph serves only as unlabeled context and diagonal self-pair edges are removed. On seven MAG benchmarks, GOMA achieves state-of-the-art or tied state-of-the-art retrieval and remains substantially more stable than the strongest graph competitor, demonstrating that MAG structure can serve as an effective post-encoder for frozen multimodal embeddings.

📄 PDF Abstract BibTeX arXiv:2605.15723

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

GoMAvatar: Efficient Animatable Human Modeling from Monocular Video Using Gaussians-on-Mesh

2024-04-11 · CVPR 2024 1 · Jing Wen, Xiaoming Zhao, Zhongzheng Ren, Alexander G. Schwing 외

We introduce GoMAvatar, a novel approach for real-time, memory-efficient, high-quality animatable human modeling. GoMAvatar takes as input a single monocular video to create a digital avatar capable of re-articulation in…

Computational Efficiency

GOMA: Proactive Embodied Cooperative Communication via Goal-Oriented Mental Alignment

2024-03-17 · Lance Ying, Kunal Jha, Shivam Aarya, Joshua B. Tenenbaum 외

Verbal communication plays a crucial role in human cooperation, particularly when the partners only have incomplete information about the task, environment, and each other's mental state. In this paper, we propose a nove…

EgoMap: Projective mapping and structured egocentric memory for Deep RL

2020-01-24 · Edward Beeching, Christian Wolf, Jilles Dibangoye, Olivier Simonin

Tasks involving localization, memorization and planning in partially observable 3D environments are an ongoing challenge in Deep Reinforcement Learning. We present EgoMap, a spatially structured neural memory architectur…

Deep Reinforcement LearningMemorizationreinforcement-learningReinforcement Learning+1

Structure-aware Knowledge-guided Heterogeneous Mamba for Zygomaticomaxillary Suture Assessment

2026-06-15 · Xiaoqi Guo, Birui Chen, Xinquan Yang, Chaoyun Zhang 외 arxiv

The Zygomaticomaxillary Suture is a key circummaxillary structure that connects the zygomatic bone and the maxilla, which serves as a primary site of resistance during maxillary advancement, and its maturation status dir…

Training-Free Color-Aware Adversarial Diffusion Sanitization for Diffusion Stegomalware Defense at Security Gateways

2025-12-30 · Vladimir Frants, Sos Agaian arxiv

The rapid expansion of generative AI has normalized large-scale synthetic media creation, enabling new forms of covert communication. Recent generative steganography methods, particularly those based on diffusion models,…