paper-with-me

홈 › Papers

Layer-Aware Video Composition via Split-then-Merge

2025-11-25 · Ozgur Kara, Yujia Chen, Ming-Hsuan Yang, James M. Rehg, Wen-Sheng Chu, Du Tran arxiv

We present Split-then-Merge (StM), a novel framework designed to enhance control in generative video composition and address its data scarcity problem. Unlike conventional methods relying on annotated datasets or handcrafted rules, StM splits a large corpus of unlabeled videos into dynamic foreground and background layers, then self-composes them to learn how dynamic subjects interact with diverse scenes. This process enables the model to learn the complex compositional dynamics required for realistic video generation. StM introduces a novel transformation-aware training pipeline that utilizes a multi-layer fusion and augmentation to achieve affordance-aware composition, alongside an identity-preservation loss that maintains foreground fidelity during blending. Experiments show StM outperforms SoTA methods in both quantitative benchmarks and in humans/VLLM-based qualitative evaluations. More details are available at our project page: https://split-then-merge.github.io

📄 PDF Abstract BibTeX arXiv:2511.20809

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

FuseFormer: Fusing Fine-Grained Information in Transformers for Video Inpainting

2021-09-07 · ICCV 2021 10 · Rui Liu, Hanming Deng, Yangyi Huang, Xiaoyu Shi 외

Transformer, as a strong and flexible architecture for modelling long-range relations, has been widely explored in vision tasks. However, when used in video inpainting that requires fine-grained representation, existed m…

Seeing Beyond the VisibleVideo Inpainting

CoordX: Accelerating Implicit Neural Representation with a Split MLP Architecture

2022-01-28 · ICLR 2022 4 · Ruofan Liang, Hongyi Sun, Nandita Vijaykumar

Implicit neural representations with multi-layer perceptrons (MLPs) have recently gained prominence for a wide variety of tasks such as novel view synthesis and 3D object representation and rendering. However, a signific…

3D Shape RepresentationNovel View Synthesis

Video-adverb retrieval with compositional adverb-action embeddings

2023-09-26 · Thomas Hummel, Otniel-Bogdan Mercea, A. Sophia Koepke, Zeynep Akata

Retrieving adverbs that describe an action in a video poses a crucial step towards fine-grained video understanding. We propose a framework for video-to-adverb retrieval (and vice versa) that aligns video embeddings with…

TripletVideo-Adverb RetrievalVideo-Adverb Retrieval (Unseen Compositions)

DynaSplit: A Hardware-Software Co-Design Framework for Energy-Aware Inference on Edge

2024-10-31 · Daniel May, Alessandro Tundo, Shashikant Ilager, Ivona Brandic

The deployment of ML models on edge devices is challenged by limited computational resources and energy availability. While split computing enables the decomposition of large neural networks (NNs) and allows partial comp…

CPUScheduling

AGQA 2.0: An Updated Benchmark for Compositional Spatio-Temporal Reasoning

2022-04-12 · Madeleine Grunde-McLaughlin, Ranjay Krishna, Maneesh Agrawala

Prior benchmarks have analyzed models' answers to questions about videos in order to measure visual compositional reasoning. Action Genome Question Answering (AGQA) is one such benchmark. AGQA provides a training/test sp…

Question Answering