paper-with-me

Papers

EffectLearner: World-Aware Object-Effect Reasoning for Real-World Video Object Removal

2026-08-06 · Feier Wu, Wanke Xia, Xu He, Zilang Zhou, Si Chen, Dongxia Liu, Liyang Chen, Qimeng Wu, Zhengbo Zhang, Wenming Yang, Zhiyong Wu arxiv

Video object removal must eliminate not only the target object but also its induced effects while maintaining high-fidelity and spatiotemporally coherent restoration. Existing methods mainly learn object-effect correspondences implicitly from predefined effect categories and fixed data distributions, limiting their generalization to complex real-world scenes involving compositional effects, spatially detached or weakly correlated effects, long-tail physical phenomena, and dynamically evolving interactions. We propose EffectLearner, a semantic-reasoning-enhanced framework that combines a VLM-based Object-Effect Reasoner with a DiT-based Video Eraser. Guided by a structured effect-analysis prompt, the Reasoner performs cross-modal reasoning over a target-highlighted video and extracts compact effect-aware context, which guides the Video Eraser toward comprehensive object-effect removal. Motion-aware mask guidance and motion-consistency supervision further improve removal coverage and spatiotemporal stability under object motion and evolving scene dynamics. To fully exploit the framework in challenging real-world scenarios, we further construct EffectWorld, a paired video dataset specifically designed for complex object-induced effects, and introduce a progressive training curriculum that combines common supervision with complex-effect data. On the standard ROSE-Bench, EffectLearner outperforms existing baselines on most metrics and achieves clear advantages on both EffectWorld-Eval and the challenging EffectWorld-Wild, demonstrating its ability to deliver high-quality video object removal in complex real-world scenes.

📄 PDF Abstract BibTeX arXiv:2608.05565

Code (3)

MorleyOlsen/EffectLearner-Official ★ 6
Tavish9/awesome-daily-AI-arxiv ★ 113
arxivsub/arXivSub_daily_arxiv ★ 4

Similar Papers 제목 키워드 기반

OFlow: Injecting Object-Aware Temporal Flow Matching for Robust Robotic Manipulation

2026-04-20 · Kuanning Wang, Ke Fan, Chenhao Qiu, Zeyu Shangguan 외 arxiv

Robust robotic manipulation requires not only predicting how the scene evolves over time, but also recognizing task-relevant objects in complex scenes. However, existing VLA models face two limitations. They typically ac…

PRIMEDrive-CoT: A Precognitive Chain-of-Thought Framework for Uncertainty-Aware Object Interaction in Driving Scene Scenario

2025-04-08 · Sriram Mandalika, Lalitha V, Athira Nambiar

Driving scene understanding is a critical real-world problem that involves interpreting and associating various elements of a driving environment, such as vehicles, pedestrians, and traffic signals. Despite advancements …

3D Object DetectionAutonomous DrivingObjectobject-detection+2

SAT: Dynamic Spatial Aptitude Training for Multimodal Language Models

2024-12-10 · Arijit Ray, Jiafei Duan, Ellis Brown, Reuben Tan 외

Reasoning about motion and space is a fundamental cognitive capability that is required by multiple real-world applications. While many studies highlight that large multimodal language models (MLMs) struggle to reason ab…

Action RecognitionSpatial Reasoning

GLOVER: Generalizable Open-Vocabulary Affordance Reasoning for Task-Oriented Grasping

2024-11-19 · Teli Ma, Zifan Wang, Jiaming Zhou, Mengmeng Wang 외

Inferring affordable (i.e., graspable) parts of arbitrary objects based on human specifications is essential for robots advancing toward open-vocabulary manipulation. Current grasp planners, however, are hindered by limi…

Common Sense ReasoningHuman-Object Interaction DetectionPose EstimationWorld Knowledge

Occlusion-Aware Search for Object Retrieval in Clutter

2020-11-06 · Wissam Bejjani, Wisdom C. Agboh, Mehmet R. Dogar, Matteo Leonetti

We address the manipulation task of retrieving a target object from a cluttered shelf. When the target object is hidden, the robot must search through the clutter for retrieving it. Solving this task requires reasoning o…

ObjectRetrieval