paper-with-me

홈 › Papers

VQQA: An Agentic Approach for Video Evaluation and Quality Improvement

2026-03-12 · Yiwen Song, Tomas Pfister, Yale Song arxiv

Despite rapid advancements in video generation models, aligning their outputs with complex user intent remains challenging. Existing test-time optimization methods are typically either computationally expensive or require white-box access to model internals. To address this, we present VQQA (Video Quality Question Answering), a unified, multi-agent framework generalizable across diverse input modalities and video generation tasks. By dynamically generating visual questions and using the resulting Vision-Language Model (VLM) critiques as semantic gradients, VQQA replaces traditional, passive evaluation metrics with human-interpretable, actionable feedback. This enables a highly efficient, closed-loop prompt optimization process via a black-box natural language interface. Extensive experiments demonstrate that VQQA effectively isolates and resolves visual artifacts, substantially improving generation quality in just a few refinement steps. Applicable to both text-to-video (T2V) and image-to-video (I2V) tasks, our method achieves absolute improvements of +11.57% on T2V-CompBench and +8.43% on VBench2 over vanilla generation, significantly outperforming state-of-the-art stochastic search and prompt optimization techniques.

📄 PDF Abstract BibTeX arXiv:2603.12310

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringVideo Generation

Similar Papers 제목 키워드 기반

COMIC: Agentic Sketch Comedy Generation

2026-03-11 · Susung Hong, Brian Curless, Ira Kemelmacher-Shlizerman, Steve Seitz arxiv

We propose a fully automated AI system that produces short comedic videos similar to sketch shows such as Saturday Night Live. Starting with character references, the system employs a population of agents loosely based o…

Video Generation

AVA-Encoder: Towards Agent-Native Video Representation Learning

2026-08-12 · Chuyue Li, Jinpeng Yu, Haozhe Wang, Tian Xueyun 외 hf

Creative agents still lack an effective way to learn from high-quality human films, limiting their ability to produce cinematic-grade videos. A key challenge is the absence of a structured video representation that is bo…

Representation LearningVideo Reconstruction

AMUSE: Audio-Visual Benchmark and Alignment Framework for Agentic Multi-Speaker Understanding

2025-12-18 · Sanjoy Chowdhury, Karren D. Yang, Xudong Liu, Fartash Faghri 외 arxiv

Recent multimodal large language models (MLLMs) such as GPT-4o and Qwen3-Omni show strong perception but struggle in multi-speaker, dialogue-centric settings that demand agentic reasoning tracking who speaks, maintaining…

A$^2$RD: Agentic Autoregressive Diffusion for Long Video Consistency

2026-05-07 · Do Xuan Long, Yale Song, Min-Yen Kan, Tomas Pfister 외 arxiv

Synthesizing consistent and coherent long video remains a fundamental challenge. Existing methods suffer from semantic drift and narrative collapse over long horizons. We present A$^2$RD, an Agentic Auto-Regressive Diffu…

VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos

2026-02-08 · Wenqi Liu, Yunxiao Wang, Shijie Ma, Meng Liu 외 arxiv

In long-video understanding, conventional uniform frame sampling often fails to capture key visual evidence, leading to degraded performance and increased hallucinations. To address this, recent agentic thinking-with-vid…

Reinforcement LearningQuestion AnsweringVideo Grounding