paper-with-me

Papers

Leveraging Weak Cross-Modal Guidance for Coherence Modelling via Iterative Learning

2024-08-01 · Yi Bin, Junrong Liao, Yujuan Ding, Haoxuan Li, Yang Yang, See-Kiong Ng, Heng Tao Shen

Cross-modal coherence modeling is essential for intelligent systems to help them organize and structure information, thereby understanding and creating content of the physical world coherently like human-beings. Previous work on cross-modal coherence modeling attempted to leverage the order information from another modality to assist the coherence recovering of the target modality. Despite of the effectiveness, labeled associated coherency information is not always available and might be costly to acquire, making the cross-modal guidance hard to leverage. To tackle this challenge, this paper explores a new way to take advantage of cross-modal guidance without gold labels on coherency, and proposes the Weak Cross-Modal Guided Ordering (WeGO) model. More specifically, it leverages high-confidence predicted pairwise order in one modality as reference information to guide the coherence modeling in another. An iterative learning paradigm is further designed to jointly optimize the coherence modeling in two modalities with selected guidance from each other. The iterative cross-modal boosting also functions in inference to further enhance coherence prediction in each modality. Experimental results on two public datasets have demonstrated that the proposed method outperforms existing methods for cross-modal coherence modeling tasks. Major technical modules have been evaluated effective through ablation studies. Codes are available at: \url{https://github.com/scvready123/IterWeGO}.

📄 PDF Abstract BibTeX arXiv:2408.00305

Code (1)

scvready123/iterwego 공식 구현 pytorch

Similar Papers 제목 키워드 기반

A Multimodal Approach Combining Structural and Cross-domain Textual Guidance for Weakly Supervised OCT Segmentation

2024-11-19 · Jiaqi Yang, Nitish Mehta, Xiaoling Hu, Chao Chen 외

Accurate segmentation of Optical Coherence Tomography (OCT) images is crucial for diagnosing and monitoring retinal diseases. However, the labor-intensive nature of pixel-level annotation limits the scalability of superv…

DescriptiveDiagnosticSegmentationSemantic Segmentation+2

Group Contrastive Learning for Weakly Paired Multimodal Data

2026-02-03 · Aditya Gorla, Hugues Van Assel, Jan-Christian Huetter, Heming Yao 외 arxiv

We present GROOVE, a semi-supervised multi-modal representation learning approach for high-content perturbation data where samples across modalities are weakly paired through shared perturbation labels but lack direct co…

Representation LearningContrastive Learning

Self-Rewarded Multimodal Coherent Reasoning Across Diverse Visual Domains

2025-12-27 · Jesen Zhang, Ningyuan Liu, Kaitong Cai, Sidi Liu 외 arxiv

Multimodal LLMs often produce fluent yet unreliable reasoning, exhibiting weak step-to-step coherence and insufficient visual grounding, largely because existing alignment approaches supervise only the final answer while…

Visual Grounding

Multimodal Procedural Planning via Dual Text-Image Prompting

2023-05-02 · Yujie Lu, Pan Lu, Zhiyu Chen, Wanrong Zhu 외

Embodied agents have achieved prominent performance in following human instructions to complete tasks. However, the potential of providing instructions informed by texts and images to assist humans in completing tasks re…

Image GenerationImage to textInformativenessText to Image Generation+1

Training-Free Multimodal Guidance for Video to Audio Generation

2025-09-29 · Eleonora Grassucci, Giuliano Galadini, Giordano Cicchetti, Aurelio Uncini 외 arxiv

Video-to-audio (V2A) generation aims to synthesize realistic and semantically aligned audio from silent videos, with potential applications in video editing, Foley sound design, and assistive multimedia. Although the exc…

Audio Generation