paper-with-me

Papers

Reframe Anything: LLM Agent for Open World Video Reframing

2024-03-10 · Jiawang Cao, Yongliang Wu, Weiheng Chi, Wenbo Zhu, Ziyue Su, Jay Wu

The proliferation of mobile devices and social media has revolutionized content dissemination, with short-form video becoming increasingly prevalent. This shift has introduced the challenge of video reframing to fit various screen aspect ratios, a process that highlights the most compelling parts of a video. Traditionally, video reframing is a manual, time-consuming task requiring professional expertise, which incurs high production costs. A potential solution is to adopt some machine learning models, such as video salient object detection, to automate the process. However, these methods often lack generalizability due to their reliance on specific training data. The advent of powerful large language models (LLMs) open new avenues for AI capabilities. Building on this, we introduce Reframe Any Video Agent (RAVA), a LLM-based agent that leverages visual foundation models and human instructions to restructure visual content for video reframing. RAVA operates in three stages: perception, where it interprets user instructions and video content; planning, where it determines aspect ratios and reframing strategies; and execution, where it invokes the editing tools to produce the final video. Our experiments validate the effectiveness of RAVA in video salient object detection and real-world reframing tasks, demonstrating its potential as a tool for AI-powered video editing.

📄 PDF Abstract BibTeX arXiv:2403.06070

Code (0)

등록된 구현이 없습니다.

Tasks

object-detectionObject DetectionSalient Object DetectionVideo EditingVideo Salient Object Detection

Similar Papers 제목 키워드 기반

Tracking Anything with Decoupled Video Segmentation

2023-09-07 · ICCV 2023 1 · Ho Kei Cheng, Seoung Wug Oh, Brian Price, Alexander Schwing 외

Training data for video segmentation are expensive to annotate. This impedes extensions of end-to-end algorithms to new video segmentation tasks, especially in large-vocabulary settings. To 'track anything' without train…

Open-Vocabulary Video SegmentationOpen-World Video SegmentationPanoptic SegmentationReferring Expression Segmentation+9

Scaling Instructable Agents Across Many Simulated Worlds

2024-03-13 · SIMA Team, Maria Abi Raad, Arun Ahuja, Catarina Barros 외

Building embodied AI systems that can follow arbitrary language instructions in any 3D environment is a key challenge for creating general AI. Accomplishing this goal requires learning to ground language in perception an…

RLAnything: Forge Environment, Policy, and Reward Model in Completely Dynamic RL System

2026-02-02 · Yinjie Wang, Tianbao Xie, Ke Shen, Mengdi Wang 외 arxiv

We propose RLAnything, a reinforcement learning framework that dynamically forges environment, policy, and reward models through closed-loop optimization, amplifying learning signals and strengthening the overall RL syst…

Reinforcement Learning

MedSAM3: Delving into Segment Anything with Medical Concepts

2025-11-24 · Anglin Liu, Rundong Xue, Xu R. Cao, Yifan Shen 외 arxiv

Medical image segmentation is fundamental for biomedical discovery. Existing methods lack generalizability and demand extensive, time-consuming manual annotation for new clinical application. Here, we propose MedSAM-3, a…

Medical Image SegmentationVideo Segmentation

SAVMap: Structure-Aided Visual Mapping of Large-Scale 2.5D Manhattan Wireframes from Panoramic Video

2026-06-01 · Howard Huang, Bharath Surianarayanan, Keifer Lee, Chenyu Wang 외 arxiv

Precise 3D representations of industrial environments enable tasks such as robot localization and digital twin generation. We propose SAVMap, a method for generating a semantic wireframe map of warehouse shelf and light …

Semantic Segmentation