paper-with-me

홈 › Papers

AVA-Encoder: Towards Agent-Native Video Representation Learning

2026-08-12 · Chuyue Li, Jinpeng Yu, Haozhe Wang, Tian Xueyun, Zhijing Zhang, Bingnan Li, Shuqi Gu, Kan Ren, Jiaming Liu, Ruihua Hua hf

Creative agents still lack an effective way to learn from high-quality human films, limiting their ability to produce cinematic-grade videos. A key challenge is the absence of a structured video representation that is both faithful to film content and directly usable for agentic reasoning and manipulation. To address the challenge, we propose the Agentic Video Auto-Encoder (AVA-Encoder), a framework for learning agent-native video representations via agentic auto-encoding. AVA-Encoder transforms a video into a knowledge graph (KG) representation and then reconstructs it back into video. Its hierarchy and state nodes store structured text, while a linked asset layer holds generated images, audio, and video. Typed edges preserve the relations between these text descriptions and assets in a form that agents can easily understand, query, and edit. The video reconstruction differences drive a textual-gradient optimization framework, which expresses evaluation feedback as natural-language update directions for Data-Independent Encoding Policy Pseudo-Training in the outer loop and optional Data-Dependent KG Representation Refinement in the test-time inner loop. Extensive experiments show that AVA-Encoder improves by 20.7 percentage points over the strongest external baseline. In the controlled policy-only setting, its pseudo-trained shot-level Agentic Video Encoder policy also outperforms a carefully human-tuned policy while using 74.3% fewer system-prompt tokens. We release the complete AVA-Encoder framework, a reliable agentic video reconstruction benchmark, and the first dataset of high-quality film KG representations.

📄 PDF Abstract BibTeX arXiv:2608.12313

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningVideo Reconstruction

Similar Papers 제목 키워드 기반

TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models

2025-12-01 · Zhiheng Liu, Weiming Ren, Haozhe Liu, Zijian Zhou 외 arxiv

Unified multimodal models (UMMs) aim to jointly perform multimodal understanding and generation within a single framework. We present TUNA, a native UMM that builds a unified continuous visual representation by cascading…

Video GenerationImage Editing

MIST: Multiple Instance Self-Training Framework for Video Anomaly Detection

2021-04-04 · CVPR 2021 1 · Jia-Chang Feng, Fa-Ting Hong, Wei-Shi Zheng

Weakly supervised video anomaly detection (WS-VAD) is to distinguish anomalies from normal events based on discriminative representations. Most existing works are limited in insufficient video representations. In this wo…

Anomaly DetectionAnomaly Detection In Surveillance VideosPseudo LabelWeakly-supervised Video Anomaly Detection

Adversarially Masked Video Consistency for Unsupervised Domain Adaptation

2024-03-24 · Xiaoyu Zhu, Junwei Liang, Po-Yao Huang, Alex Hauptmann

We study the problem of unsupervised domain adaptation for egocentric videos. We propose a transformer-based model to learn class-discriminative and domain-invariant feature representations. It consists of two novel desi…

Domain AdaptationUnsupervised Domain Adaptation

MED-VT++: Unifying Multimodal Learning with a Multiscale Encoder-Decoder Video Transformer

2023-04-12 · CVPR 2023 1 · Rezaul Karim, He Zhao, Richard P. Wildes, Mennatullah Siam

In this paper, we present an end-to-end trainable unified multiscale encoder-decoder transformer that is focused on dense prediction tasks in video. The presented Multiscale Encoder-Decoder Video Transformer (MED-VT) use…

Action SegmentationDecoderOptical Flow EstimationSegmentation+5

MAVIS: Multi-Agent Video Retrieval via Structured Video Understanding

2026-06-08 · Jie Zhang, Qilang Ye, Hao Zhou, Haochen Liang 외 arxiv

The dominant paradigm in video retrieval relies on embedding-based full-corpus scanning, which suffers from inherent computational inefficiency and the semantic asymmetry between information-dense videos and sparse textu…

Video Retrieval