paper-with-me

홈 › Papers

All modalities are equal, but video is more equal: Closing the Cross-Attention Gap in Joint Video Generation

2026-09-23 · Ohad Rahamim, Dvir Samuel, Idan Schwartz, Gal Chechik hf

Video is a rich representation of a physical event, capturing appearance, geometry, motion, and temporal evolution. Other modalities, such as 3D body motion or audio, encode narrower aspects of the same event. We find that joint multimodal diffusion transformers exhibit a corresponding asymmetry in cross-modal correspondence: companion modalities develop strong correspondences to video, but the reciprocal correspondences through which they constrain video remain substantially weaker. We express both directions as comparable correspondence distributions over video tokens and define their disagreement as the reciprocal correspondence gap. We introduce RecCAR, standing for Reciprocal Cross-modal Attention Regularization, a KL regularizer that uses the well-established video-to-modality correspondence as a fixed reference and aligns the weaker modality-to-video correspondence toward it. Across joint video-motion and video-audio generation, RecCAR improves the Human Anatomy score from 0.69 to 0.75 and reduces audio-video desynchronization from 0.804 to 0.752, while improving overall generation

📄 PDF Abstract BibTeX arXiv:2609.27901

Code (3)

Tavish9/awesome-daily-AI-arxiv ★ 121
Valiant-Cat/hfpaper
will-rice/tts-papers ★ 4

Tasks

Audio GenerationVideo Generation

Similar Papers 제목 키워드 기반

Closing the U.S. gender wage gap requires understanding its heterogeneity

2018-12-11 · Philipp Bach, Victor Chernozhukov, Martin Spindler

In 2016, the majority of full-time employed women in the U.S. earned significantly less than comparable men. The extent to which women were affected by gender inequality in earnings, however, depended greatly on socio-ec…

regression

Revisiting Non-separable Binary Classification and its Applications in Anomaly Detection

2023-12-03 · Matthew Lau, Ismaila Seck, Athanasios P Meliopoulos, Wenke Lee 외

The inability to linearly classify XOR has motivated much of deep learning. We revisit this age-old problem and show that linear classification of XOR is indeed possible. Instead of separating data between halfspaces, we…

Anomaly DetectionBinary ClassificationSupervised Anomaly Detection

Event-Equalized Dense Video Captioning

2025-01-01 · CVPR 2025 1 · Kangyi Wu, Pengna Li, Jingwen Fu, Yizhe Li 외

Dense video captioning aims to localize and caption all events in arbitrary untrimmed videos. Although previous methods have achieved appealing results, they still face the issue of temporal bias, i.e, models tend to…

Dense Video CaptioningVideo Captioning

A Survey on Multimodal Disinformation Detection

2021-03-13 · COLING 2022 10 · Firoj Alam, Stefano Cresci, Tanmoy Chakraborty, Fabrizio Silvestri 외

Recent years have witnessed the proliferation of offensive content online such as fake news, propaganda, misinformation, and disinformation. While initially this was mostly about textual content, over time images and vid…

MisinformationSurvey

CFSum: A Transformer-Based Multi-Modal Video Summarization Framework With Coarse-Fine Fusion

2025-03-01 · Yaowei Guo, Jiazheng Xing, Xiaojun Hou, Shuo Xin 외

Video summarization, by selecting the most informative and/or user-relevant parts of original videos to create concise summary videos, has high research value and consumer demand in today's video proliferation era. Multi…

Video Summarization