paper-with-me

홈 › Papers

MoVA: Learning Asymmetric Dual Projections for Modular Long Video-Text Alignment

2026-07-01 · Peiyuan Zhu, Shaoan Xie, Zijian Li, Yifan Shen, Namrata Deka, Harsh Shrivastava, Guangyi Chen, Kun Zhang arxiv

Contrastive pre-training has propelled video-text alignment, yet models often inherit the critical limitations of their image-text predecessors like CLIP, resulting in entangled representations. These challenges are severely exacerbated by two fundamental properties in the video domain: Temporal Misalignment, where textual descriptions often correlate only to specific, constrained temporal windows, leaving other frames text-irrelevant; and Semantic Asymmetry, which dictates a sparse, bidirectional, and non-equivalent relevance between frame-level visual details and caption-level concepts. This failure persists whether captions are short and temporally disjoint, creating ambiguity, or long and detailed, fostering entanglement between static objects and their temporal evolution. In this paper, we establish theoretical conditions that enable flexible alignment between video and text representations across the temporal dimension and at varying levels of granularity. Building on these theoretical insights, we introduce MoVA, Modular Long Video-Text Alignment, which learns dual asymmetric projections: a text-side projection that adaptively selects frame-aware subspaces of the caption, and a video-side projection that disentangles text-relevant visual concepts. Our framework ensures that the model can preserve global cross-modal semantics while disentangling evolving, frame-specific concepts and scale naturally to long captions and videos. Empirical evaluations show that MoVA outperforms existing methods in multiple video-text alignment tasks, demonstrating the effectiveness of our method.

📄 PDF Abstract BibTeX arXiv:2607.00858

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Asymmetric Dual-Decoder U-Net for Joint Rain and Haze Removal

2022-06-14 · Yuan Feng, Yaojun Hu, Pengfei Fang, Yanhong Yang 외

This work studies the joint rain and haze removal problem. In real-life scenarios, rain and haze, two often co-occurring common weather phenomena, can greatly degrade the clarity and quality of the scene images, leading …

Autonomous DrivingDecoderSingle Particle Analysis

CXR-LT 2026 Challenge: Projection-Aware Multi-Label and Zero-Shot Chest X-Ray Classification

2026-04-02 · Juno Cho, Dohui Kim, Mingeon Kim, Hyunseo Jang 외 arxiv

This challenge tackles multi-label classification for known chest X-ray (CXR) lesions and zero-shot classification for unseen ones. To handle diverse CXR projections, we integrate projection-specific models via a classif…

Multi-Label ClassificationZero-shot GeneralizationContrastive Learning

Asymmetric Feature Maps with Application to Sketch Based Retrieval

2017-04-12 · CVPR 2017 7 · Giorgos Tolias, Ondřej Chum

We propose a novel concept of asymmetric feature maps (AFM), which allows to evaluate multiple kernels between a query and database entries without increasing the memory requirements. To demonstrate the advantages of the…

Image RetrievalRetrievalSketch-Based Image RetrievalTranslation

Reusing Combinatorial Structure: Faster Iterative Projections over Submodular Base Polytopes

2021-06-22 · NeurIPS 2021 12 · Jai Moondra, Hassan Mortagy, Swati Gupta

Optimization algorithms such as projected Newton's method, FISTA, mirror descent, and its variants enjoy near-optimal regret bounds and convergence rates, but suffer from a computational bottleneck of computing ``project…

Quadratic Decomposable Submodular Function Minimization

2018-06-26 · NeurIPS 2018 12 · Pan Li, Niao He, Olgica Milenkovic

We introduce a new convex optimization problem, termed quadratic decomposable submodular function minimization. The problem is closely related to decomposable submodular function minimization and arises in many learning …