paper-with-me

홈 › Papers

See More, Store Less: Memory-Efficient Resolution for Video Moment Retrieval

2026-01-14 · Mingyu Jeon, Sungjin Han, Jinkwon Hwang, Minchol Kwon, Jonghee Kim, Junyeong Kim arxiv

Recent advances in Multimodal Large Language Models (MLLMs) have improved image recognition and reasoning, but video-related tasks remain challenging due to memory constraints from dense frame processing. Existing Video Moment Retrieval (VMR) methodologies rely on sparse frame sampling, risking potential information loss, especially in lengthy videos. We propose SMORE (See MORE, store less), a framework that enhances memory efficiency while maintaining high information resolution. SMORE (1) uses query-guided captions to encode semantics aligned with user intent, (2) applies query-aware importance modulation to highlight relevant segments, and (3) adaptively compresses frames to preserve key content while reducing redundancy. This enables efficient video understanding without exceeding memory budgets. Experimental validation reveals that SMORE achieves state-of-the-art performance on QVHighlights, Charades-STA, and ActivityNet-Captions benchmarks.

📄 PDF Abstract BibTeX arXiv:2601.09350

Code (0)

등록된 구현이 없습니다.

Tasks

Moment Retrieval

Similar Papers 제목 키워드 기반

LiVOS: Light Video Object Segmentation with Gated Linear Matching

2024-11-05 · CVPR 2025 1 · Qin Liu, JianFeng Wang, Zhengyuan Yang, Linjie Li 외

Semi-supervised video object segmentation (VOS) has been largely driven by space-time memory (STM) networks, which store past frame features in a spatiotemporal memory to segment the current frame via softmax attention. …

GPUSemantic SegmentationSemi-Supervised Video Object SegmentationVideo Object Segmentation+1

XMem: Long-Term Video Object Segmentation with an Atkinson-Shiffrin Memory Model

2022-07-14 · Ho Kei Cheng, Alexander G. Schwing

We present XMem, a video object segmentation architecture for long videos with unified feature memory stores inspired by the Atkinson-Shiffrin memory model. Prior work on video object segmentation typically only uses one…

2D Human Pose Estimation2D Object Detection3D Absolute Human Pose EstimationSegmentation+4

Vista: Scene-Aware Optimization for Streaming Video Question Answering under Post-Hoc Queries

2026-02-09 · Haocheng Lu, Nan Zhang, Wei Tao, Xiaoyang Qu 외 arxiv

Streaming video question answering (Streaming Video QA) poses distinct challenges for multimodal large language models (MLLMs), as video frames arrive sequentially and user queries can be issued at arbitrary time points.…

Video Question Answering

Information Prebuilt Recurrent Reconstruction Network for Video Super-Resolution

2021-12-10 · Shuyun Wang, Ming Yu, Cuihong Xue, Yingchun Guo 외

The video super-resolution (VSR) method based on the recurrent convolutional network has strong temporal modeling capability for video sequences. However, the temporal receptive field of different recurrent units in the …

Super-ResolutionVideo Super-Resolution

MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

2024-04-08 · CVPR 2024 1 · Bo He, Hengduo Li, Young Kyun Jang, Menglin Jia 외

With the success of large language models (LLMs), integrating the vision model into LLMs to build vision-language foundation models has gained much more interest recently. However, existing LLM-based large multimodal mod…

GPUMultiple-choiceQuestion AnsweringTemporal Relation Extraction+5