Papers Video Editing
“Video Editing” 태그가 달린 논문 346편 · 필터 해제
DFVEdit: Conditional Delta Flow Vector for Zero-shot Video Editing
The advent of Video Diffusion Transformers (Video DiTs) marks a milestone in video generation. However, directly applying existing video editing methods to Video DiTs often incurs substantial computational overhead, due …
Video EditingVideo GenerationLet Your Video Listen to Your Music!
Aligning the rhythm of visual motion in a video with a given music track is a practical need in multimedia production, yet remains an underexplored task in autonomous video editing. Effective alignment between motion and…
GPUMusic GenerationRhythmVideo Editing+1Causally Steered Diffusion for Automated Video Counterfactual Generation
Adapting text-to-image (T2I) latent diffusion models for video editing has shown strong visual fidelity and controllability, but challenges remain in maintaining causal relationships in video content. Edits affecting cau…
counterfactualVideo EditingVideo GenerationLoRA-Edit: Controllable First-Frame-Guided Video Editing via Mask-Aware LoRA Fine-Tuning
Video editing using diffusion models has achieved remarkable results in generating high-quality edits for videos. However, current methods often rely on large-scale pretraining, limiting flexibility for specific edits. F…
Video EditingRoboSwap: A GAN-driven Video Diffusion Framework For Unsupervised Robot Arm Swapping
Recent advancements in generative models have revolutionized video synthesis and editing. However, the scarcity of diverse, high-quality datasets continues to hinder video-conditioned robotic learning, limiting cross-pla…
Video EditingSuper Encoding Network: Recursive Association of Multi-Modal Encoders for Video Understanding
Video understanding has been considered as one critical step towards world modeling, which is an important long-term problem in AI research. Recently, multi-modal foundation models have shown such potential via large-sca…
Contrastive LearningVideo EditingVideo UnderstandingFADE: Frequency-Aware Diffusion Model Factorization for Video Editing
Recent advancements in diffusion frameworks have significantly enhanced video editing, achieving high fidelity and strong alignment with textual prompts. However, conventional approaches using image diffusion models fall…
Video EditingFlowDirector: Training-Free Flow Steering for Precise Text-to-Video Editing
Text-driven video editing aims to modify video content according to natural language instructions. While recent training-free approaches have made progress by leveraging pre-trained diffusion models, they typically rely …
Text-to-Video EditingVideo EditingFullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers
Fine-grained and efficient controllability on video diffusion transformers has raised increasing desires for the applicability. Recently, In-context Conditioning emerged as a powerful paradigm for unified conditional vid…
Video EditingVideo GenerationMiniMax-Remover: Taming Bad Noise Helps Video Object Removal
Recent advances in video diffusion models have driven rapid progress in video editing techniques. However, video object removal, a critical subtask of video editing, remains challenging due to issues such as hallucinated…
Video EditingVideo GenerationZero-to-Hero: Zero-Shot Initialization Empowering Reference-Based Video Appearance Editing
Appearance editing according to user needs is a pivotal task in video editing. Existing text-guided methods often lead to ambiguities regarding user intentions and restrict fine-grained control over editing specific aspe…
Optical Flow EstimationVideo EditingVideo RestorationVideo Editing for Audio-Visual Dubbing
Visual dubbing, the synchronization of facial movements with new speech, is crucial for making content accessible across different languages, enabling broader global reach. However, current methods face significant limit…
Video EditingTDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs
Text-driven video editing is rapidly advancing, yet its rigorous evaluation remains challenging due to the absence of dedicated video quality assessment (VQA) models capable of discerning the nuances of editing quality. …
BenchmarkingLarge Language ModelVideo EditingVideo Quality Assessment+1SRDiffusion: Accelerate Video Diffusion Inference via Sketching-Rendering Cooperation
Leveraging the diffusion transformer (DiT) architecture, models like Sora, CogVideoX and Wan have achieved remarkable progress in text-to-video, image-to-video, and video editing tasks. Despite these advances, diffusion-…
Video EditingVideo GenerationREGen: Multimodal Retrieval-Embedded Generation for Long-to-Short Video Editing
Short videos are an effective tool for promoting contents and improving knowledge accessibility. While existing extractive video summarization methods struggle to produce a coherent narrative, existing abstractive method…
Language ModelingLanguage ModellingLarge Language ModelRetrieval+2From Shots to Stories: LLM-Assisted Video Editing with Unified Language Representations
Large Language Models (LLMs) and Vision-Language Models (VLMs) have demonstrated remarkable reasoning and generalization capabilities in video understanding; however, their application in video editing remains largely un…
Video EditingVideo UnderstandingDAPE: Dual-Stage Parameter-Efficient Fine-Tuning for Consistent Video Editing with Diffusion Models
Video generation based on diffusion models presents a challenging multimodal task, with video editing emerging as a pivotal direction in this field. Recent video editing approaches primarily fall into two categories: tra…
parameter-efficient fine-tuningVideo AlignmentVideo EditingVideo GenerationVideo Forgery Detection for Surveillance Cameras: A Review
The widespread availability of video recording through smartphones and digital devices has made video-based evidence more accessible than ever. Surveillance footage plays a crucial role in security, law enforcement, and …
Frame Duplication DetectionMisinformationVideo EditingA Rusty Link in the AI Supply Chain: Detecting Evil Configurations in Model Repositories
Recent advancements in large language models (LLMs) have spurred the development of diverse AI applications from code generation and video editing to text generation; however, AI supply chains such as Hugging Face, which…
Code GenerationText GenerationVideo EditingControllable Weather Synthesis and Removal with Video Diffusion Models
Generating realistic and controllable weather effects in videos is valuable for many applications. Physics-based weather simulation requires precise reconstructions that are hard to scale to in-the-wild videos, while cur…
Video Editing