paper-with-me

Papers

OmniVid: A Generative Framework for Universal Video Understanding

2024-03-26 · CVPR 2024 1 · Junke Wang, Dongdong Chen, Chong Luo, Bo He, Lu Yuan, Zuxuan Wu, Yu-Gang Jiang

The core of video understanding tasks, such as recognition, captioning, and tracking, is to automatically detect objects or actions in a video and analyze their temporal evolution. Despite sharing a common goal, different tasks often rely on distinct model architectures and annotation formats. In contrast, natural language processing benefits from a unified output space, i.e., text sequences, which simplifies the training of powerful foundational language models, such as GPT-3, with extensive training corpora. Inspired by this, we seek to unify the output space of video understanding tasks by using languages as labels and additionally introducing time and box tokens. In this way, a variety of video tasks could be formulated as video-grounded token generation. This enables us to address various types of video tasks, including classification (such as action recognition), captioning (covering clip captioning, video question answering, and dense video captioning), and localization tasks (such as visual object tracking) within a fully shared encoder-decoder architecture, following a generative framework. Through comprehensive experiments, we demonstrate such a simple and straightforward idea is quite effective and can achieve state-of-the-art or competitive results on seven video benchmarks, providing a novel perspective for more universal video understanding. Code is available at https://github.com/wangjk666/OmniVid.

📄 PDF Abstract BibTeX arXiv:2403.17935

Code (1)

wangjk666/omnivid 공식 구현 pytorch

Tasks

Action RecognitionDecoderDense Video CaptioningObject TrackingQuestion AnsweringVideo CaptioningVideo Question AnsweringVideo UnderstandingVisual Object Tracking

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
{Dispute@FaQ-s}How to file a dispute with Expedia? How to file a dispute with Expedia? To file a complaint against Expedia, first try contacting their customer service directly. You can reach them by phone at…
15 Ways to Contact How can i speak to someone at Delta Airlines 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…
Residual Connection 설명 없음
Cosine Annealing Cosine Annealing is a type of learning rate schedule that has the effect of starting with a large learning rate that is relatively rapidly decreased to a minimum value before…

Similar Papers 제목 키워드 기반

OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention

2026-02-05 · Zhangquan Chen, Jiale Tao, Ruihuang Li, Yihao Hu 외 arxiv

While humans perceive the world through diverse modalities that operate synergistically to support a holistic understanding of their surroundings, existing omnivideo models still face substantial challenges on audio-visu…

Self-Supervised LearningContrastive LearningVisual Reasoning

OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs

2025-10-12 · Caorui Li, Yu Chen, Yiyan Ji, Jin Xu 외 arxiv

Recent advances in multimodal large language models (MLLMs) have demonstrated substantial potential in video understanding. However, existing benchmarks fail to comprehensively evaluate synergistic reasoning capabilities…

Causal InferenceVisual Reasoning

OmniVideo-100K: A Dataset for Audio-Visual Reasoning through Structured Scripts and Evidence Chains

2026-06-12 · Xinyue Cai, Chaoyou Fu, Yi-Fan Zhang, Ran He 외 arxiv

Current automated pipelines for audio-visual Question Answering (QA) generally adopt a ``video-caption-QA'' paradigm. However, these methods typically segment videos into short clips and generate separate descriptions fo…

Audio-visual Question AnsweringVisual Reasoning

AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression

2026-06-23 · Yijing Chen, Wenhui Tan, Xiaoyi Yu, Yuyue Wang 외 arxiv

Multimodal Large Language Models have achieved remarkable progress in short-form audio-video understanding, yet long-form audio-video comprehension remains challenged by limited context windows and severe information red…

Information Retrieval

OmniVidar: Omnidirectional Depth Estimation From Multi-Fisheye Images

2023-01-01 · CVPR 2023 1 · Sheng Xie, Daochuan Wang, Yun-hui Liu

Estimating depth from four large field of view (FoV) cameras has been a difficult and understudied problem. In this paper, we proposed a novel and simple system that can convert this difficult problem into easier bin…

Depth Estimation