paper-with-me

Papers

Agentic Aerial Cinematography: From Dialogue Cues to Cinematic Trajectories

2025-09-19 · Yifan Lin, Sophie Ziyu Liu, Ran Qi, George Z. Xue, Xinping Song, Chao Qin, Hugh H. -T. Liu arxiv

We present Agentic Aerial Cinematography: From Dialogue Cues to Cinematic Trajectories (ACDC), an autonomous drone cinematography system driven by natural language communication between human directors and drones. The main limitation of previous drone cinematography workflows is that they require manual selection of waypoints and view angles based on predefined human intent, which is labor-intensive and yields inconsistent performance. In this paper, we propose employing large language models (LLMs) and vision foundation models (VFMs) to convert free-form natural language prompts directly into executable indoor UAV video tours. Specifically, our method comprises a vision-language retrieval pipeline for initial waypoint selection, a preference-based Bayesian optimization framework that refines poses using aesthetic feedback, and a motion planner that generates safe quadrotor trajectories. We validate ACDC through both simulation and hardware-in-the-loop experiments, demonstrating that it robustly produces professional-quality footage across diverse indoor scenes without requiring expertise in robotics or cinematography. These results highlight the potential of embodied AI agents to close the loop from open-vocabulary dialogue to real-world autonomous aerial cinematography.

📄 PDF Abstract BibTeX arXiv:2509.16176

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Camera Artist: A Multi-Agent Framework for Cinematic Language Storytelling Video Generation

2026-04-10 · Haobo Hu, Qi Mao, Yuanhang Li, Libiao Jin arxiv

We propose Camera Artist, a multi-agent framework that models a real-world filmmaking workflow to generate narrative videos with explicit cinematic language. While recent multi-agent systems have made substantial progres…

Video Generation

ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models

2025-06-26 · Hongbo Liu, Jingwen He, Yi Jin, Dian Zheng 외

Cinematography, the fundamental visual language of film, is essential for conveying narrative, emotion, and aesthetic quality. While recent Vision-Language Models (VLMs) demonstrate strong general visual understanding, t…

Spatial ReasoningVideo Generation

Can video generation replace cinematographers? Research on the cinematic language of generated video

2024-12-16 · Xiaozhe Li, Kai Wu, Siyi Yang, YiZhan Qu 외

Recent advancements in text-to-video (T2V) generation have leveraged diffusion models to enhance visual coherence in videos synthesized from textual descriptions. However, existing research primarily focuses on object mo…

Video Generation

The Script is All You Need: An Agentic Framework for Long-Horizon Dialogue-to-Cinematic Video Generation

2026-01-25 · Chenyu Mu, Xin He, Qu Yang, Wanshun Chen 외 arxiv

Recent advances in video generation have produced models capable of synthesizing stunning visual content from simple text prompts. However, these models struggle to generate long-form, coherent narratives from high-level…

Video Generation

DreamCinema: Cinematic Transfer with Free Camera and 3D Character

2024-08-22 · Weiliang Chen, Fangfu Liu, Diankun Wu, Haowen Sun 외

We are living in a flourishing era of digital media, where everyone has the potential to become a personal filmmaker. Current research on cinematic transfer empowers filmmakers to reproduce and manipulate the visual elem…