paper-with-me

Papers

VideoPerceiver: Enhancing Fine-Grained Temporal Perception in Video Multimodal Large Language Models

2025-11-24 · Fufangchen Zhao, Liao Zhang, Daiqi Shi, Yuanjun Gao, Chen Ye, Yang Cai, Jian Gao, Danfeng Yan arxiv

We propose VideoPerceiver, a novel video multimodal large language model (VMLLM) that enhances fine-grained perception in video understanding, addressing VMLLMs' limited ability to reason about brief actions in short clips or rare transient events in long videos. VideoPerceiver adopts a two-stage training framework. During supervised fine-tuning (SFT), we construct "key-information-missing" videos by extracting event-action keywords from captions, identifying corresponding key frames, and replacing them with adjacent frames. We jointly encode original and modified video tokens with text tokens, aligning intermediate visual representations with keywords via an auxiliary contrastive loss to enhance sensitivity to fine-grained motion cues. In reinforcement learning (RL), both video variants are fed into the model to generate descriptions, and a novel relative reward ensures responses from complete videos outperform those from degraded inputs, explicitly training the model to recover temporally precise action details. We also curate a dataset of 80,000 videos with fine-grained actions and transient events. Experiments show VideoPerceiver substantially outperforms state-of-the-art VMLLMs on fine-grained action understanding and rare event captioning benchmarks, while maintaining strong performance on standard tasks. By prioritizing task-relevant visual features, our work redefines video-language model training for fine-grained perception.

📄 PDF Abstract BibTeX arXiv:2511.18823

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningAction Understanding

Similar Papers 제목 키워드 기반

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences

2026-06-26 · Yankai Yang, Yancheng Long, Bin Wen, Fan Yang 외 arxiv

Video multimodal large language models have made strong progress on open-ended video understanding, but they still lack precise local spatiotemporal perception. When two videos share almost the same global semantics and …

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection

2026-05-31 · Xin Dong, Wenjia Geng, Wenfeng Deng, Yansong Tang arxiv

Video Moment Retrieval (MR) and Highlight Detection (HD) are crucial tasks in video analysis that aim to localize specific moments and estimate clip-wise relevance based on a given text query. Recent approaches treat the…

Representation LearningHighlight DetectionMoment RetrievalVideo Grounding

Multi-Grained Feature Pruning for Video-Based Human Pose Estimation

2025-03-07 · Zhigang Wang, Shaojing Fan, Zhenguang Liu, Zheqi Wu 외

Human pose estimation, with its broad applications in action recognition and motion capture, has experienced significant advancements. However, current Transformer-based methods for video pose estimation often face chall…

Action RecognitionComputational EfficiencyPose Estimation

Dr.V: A Hierarchical Perception-Temporal-Cognition Framework to Diagnose Video Hallucination by Fine-grained Spatial-Temporal Grounding

2025-09-15 · Meng Luo, Shengqiong Wu, Liqiang Jing, Tianjie Ju 외 arxiv

Recent advancements in large video models (LVMs) have significantly enhance video understanding. However, these models continue to suffer from hallucinations, producing content that conflicts with input videos. To addres…

Towards Fine-grained Temporal Perception: Post-Training Large Audio-Language Models with Audio-Side Time Prompt

2026-04-15 · Yanfeng Shi, Pengfei Cai, Jun Liu, Qing Gu 외 arxiv

Large Audio-Language Models (LALMs) enable general audio understanding and demonstrate remarkable performance across various audio tasks. However, these models still face challenges in temporal perception (e.g., inferrin…

Reinforcement LearningSound Event DetectionAudio captioning