paper-with-me

Papers

CMMD: Contrastive Multi-Modal Diffusion for Video-Audio Conditional Modeling

2023-12-08 · Ruihan Yang, Hannes Gamper, Sebastian Braun

We introduce a multi-modal diffusion model tailored for the bi-directional conditional generation of video and audio. We propose a joint contrastive training loss to improve the synchronization between visual and auditory occurrences. We present experiments on two datasets to evaluate the efficacy of our proposed model. The assessment of generation quality and alignment performance is carried out from various angles, encompassing both objective and subjective metrics. Our findings demonstrate that the proposed model outperforms the baseline in terms of quality and generation speed through introduction of our novel cross-modal easy fusion architectural block. Furthermore, the incorporation of the contrastive loss results in improvements in audio-visual alignment, particularly in the high-correlation video-to-audio generation task.

📄 PDF Abstract BibTeX arXiv:2312.05412

Code (0)

등록된 구현이 없습니다.

Tasks

Audio Generation

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Bridging Contrastive Learning and Domain Adaptation: Theoretical Perspective and Practical Application

2025-01-28 · Gonzalo Iñaki Quintana, Laurence Vancamberg, Vincent Jugnon, Agnès Desolneux 외

This work studies the relationship between Contrastive Learning and Domain Adaptation from a theoretical perspective. The two standard contrastive losses, NT-Xent loss (Self-supervised) and Supervised Contrastive loss, a…

Contrastive LearningDomain Adaptation

GMS-CAVP: Improving Audio-Video Correspondence with Multi-Scale Contrastive and Generative Pretraining

2026-01-27 · Shentong Mo, Zehua Chen, Jun Zhu arxiv

Recent advances in video-audio (V-A) understanding and generation have increasingly relied on joint V-A embeddings, which serve as the foundation for tasks such as cross-modal retrieval and generation. While prior method…

Cross-Modal RetrievalContrastive Learning

DiffGAP: A Lightweight Diffusion Module in Contrastive Space for Bridging Cross-Model Gap

2025-03-15 · Shentong Mo, Zehua Chen, Fan Bao, Jun Zhu

Recent works in cross-modal understanding and generation, notably through models like CLAP (Contrastive Language-Audio Pretraining) and CAVP (Contrastive Audio-Visual Pretraining), have significantly enhanced the alignme…

AudioCapsAudio GenerationDenoising

Multimodal Generative Models for Bankruptcy Prediction Using Textual Data

2022-10-26 · Rogelio A. Mancisidor, Kjersti Aas

Textual data from financial filings, e.g., the Management's Discussion & Analysis (MDA) section in Form 10-K, has been used to improve the prediction accuracy of bankruptcy models. In practice, however, we cannot obtain …

AllPrediction

Measuring Differences between Conditional Distributions using Kernel Embeddings

2026-05-04 · Peter Moskvichev, Siu Lun Chau, Dino Sejdinovic arxiv

Comparing conditional distributions is a fundamental challenge in statistics and machine learning, with applications across a wide range of domains. While proposed methods for measuring discrepancies using kernel embeddi…