paper-with-me

홈 › Papers

VISTA: A Test-Time Self-Improving Video Generation Agent

2025-10-17 · Do Xuan Long, Xingchen Wan, Hootan Nakhost, Chen-Yu Lee, Tomas Pfister, Sercan Ö. Arık arxiv

Despite rapid advances in text-to-video synthesis, generated video quality remains critically dependent on precise user prompts. Existing test-time optimization methods, successful in other domains, struggle with the multi-faceted nature of video. In this work, we introduce VISTA (Video Iterative Self-improvemenT Agent), a novel multi-agent system that autonomously improves video generation through refining prompts in an iterative loop. VISTA first decomposes a user idea into a structured temporal plan. After generation, the best video is identified through a robust pairwise tournament. This winning video is then critiqued by a trio of specialized agents focusing on visual, audio, and contextual fidelity. Finally, a reasoning agent synthesizes this feedback to introspectively rewrite and enhance the prompt for the next generation cycle. Experiments on single- and multi-scene video generation scenarios show that while prior methods yield inconsistent gains, VISTA consistently improves video quality and alignment with user intent, achieving up to 60% pairwise win rate against state-of-the-art baselines. Human evaluators concur, preferring VISTA outputs in 66.4% of comparisons.

📄 PDF Abstract BibTeX arXiv:2510.15831

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

VISTA: A Controllable Platform for Generating and Auditing Egocentric Assistance Scenarios

2026-05-11 · Yu-Hsiang Liu, Yu-Chien Tang, An-Zi Yen arxiv

Evaluating whether AI agents can proactively assist humans in daily activities, ranging from routine household tasks to urgent safety-critical situations, requires diverse visual data. However, collecting such scenarios …

VistaGEN: Consistent Driving Video Generation with Fine-Grained Control Using Multiview Visual-Language Reasoning

2026-03-30 · Li-Heng Chen, Ke Cheng, Yahui Liu, Lei Shi 외 arxiv

Driving video generation has achieved much progress in controllability, video resolution, and length, but fails to support fine-grained object-level controllability for diverse driving videos, while preserving the spatio…

Video Generation

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models

2025-04-17 · Haojian Huang, Haodong Chen, Shengqiong Wu, Meng Luo 외

Large Video Models (LVMs) built upon Large Language Models (LLMs) have shown promise in video understanding but often suffer from misalignment with human intuition and video hallucination issues. To address these challen…

HallucinationVideo Understanding

VistaFormer: Scalable Vision Transformers for Satellite Image Time Series Segmentation

2024-09-13 · Ezra MacDonald, Derek Jacoby, Yvonne Coady

We introduce VistaFormer, a lightweight Transformer-based model architecture for the semantic segmentation of remote-sensing images. This model uses a multi-scale Transformer-based encoder with a lightweight decoder that…

DecoderSemantic SegmentationTime Series

Vista: Scene-Aware Optimization for Streaming Video Question Answering under Post-Hoc Queries

2026-02-09 · Haocheng Lu, Nan Zhang, Wei Tao, Xiaoyang Qu 외 arxiv

Streaming video question answering (Streaming Video QA) poses distinct challenges for multimodal large language models (MLLMs), as video frames arrive sequentially and user queries can be issued at arbitrary time points.…

Video Question Answering