paper-with-me

홈 › Papers

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding

2025-08-03 · Zuhao Yang, Yingchen Yu, Yunqing Zhao, Shijian Lu, Song Bai arxiv

Video Temporal Grounding (VTG) aims to precisely identify video event segments in response to textual queries. The outputs of VTG tasks manifest as sequences of events, each defined by precise timestamps, saliency scores, and textual descriptions. Despite recent advances, a fundamental limitation persists in existing Video Large Language Models (Video-LLMs): they process all task tokens through identical and static pathways, failing to recognize that temporal localization, saliency assessment, and textual generation represent fundamentally distinct tasks requiring specialized processing. To address this, we introduce TimeExpert, a Mixture-of-Experts (MoE)-based Video-LLM that effectively decomposes VTG tasks by dynamically routing task-specific tokens (e.g., timestamps, saliency scores) to specialized experts, with increased computational efficiency. Our design choices enable precise handling of each subtask, leading to improved event modeling across diverse VTG applications. Extensive experiments demonstrate that TimeExpert consistently achieves state-of-the-art performance on various VTG tasks such as Dense Video Captioning, Moment Retrieval, and Video Highlight Detection.

📄 PDF Abstract BibTeX arXiv:2508.01699

Code (0)

등록된 구현이 없습니다.

Tasks

Computational EfficiencyDense Video CaptioningHighlight DetectionMoment Retrieval

Similar Papers 제목 키워드 기반

TimeExpert: Boosting Long Time Series Forecasting with Temporal Mix of Experts

2025-09-27 · Xiaowen Ma, Shuning Ge, Fan Yang, Xiangyu Li 외 arxiv

Transformer-based architectures dominate time series modeling by enabling global attention over all timestamps, yet their rigid 'one-size-fits-all' context aggregation fails to address two critical challenges in real-wor…

Time Series Forecasting

Mixture of Experts Guided by Gaussian Splatters Matters: A new Approach to Weakly-Supervised Video Anomaly Detection

2025-08-08 · Giacomo D'Amicantonio, Snehashis Majhi, Quan Kong, Lorenzo Garattoni 외 arxiv

Video Anomaly Detection (VAD) is a challenging task due to the variability of anomalous events and the limited availability of labeled data. Under the Weakly-Supervised VAD (WSVAD) paradigm, only video-level labels are p…

Weakly-supervised Video Anomaly Detection

Exposing and Defending the Achilles' Heel of Video Mixture-of-Experts

2026-02-01 · Songping Wang, Qinglong Liu, Yueming Lyu, Ning Li 외 arxiv

Mixture-of-Experts (MoE) has demonstrated strong performance in video understanding tasks, yet its adversarial robustness remains underexplored. Existing attack methods often treat MoE as a unified architecture, overlook…

Adversarial Robustness

Show Me: Unifying Instructional Image and Video Generation with Diffusion Models

2025-11-21 · Yujiang Pu, Zhanbo Huang, Vishnu Boddeti, Yu Kong arxiv

Generating visual instructions in a given context is essential for developing interactive world simulators. While prior works address this problem through either text-guided image manipulation or video prediction, these …

Image ManipulationVideo GenerationVideo Prediction

VideoExpert: Augmented LLM for Temporal-Sensitive Video Understanding

2025-04-10 · Henghao Zhao, Ge-Peng Ji, Rui Yan, Huan Xiong 외

The core challenge in video understanding lies in perceiving dynamic content changes over time. However, multimodal large language models struggle with temporal-sensitive video tasks, which requires generating timestamps…

Instruction FollowingVideo Understanding