paper-with-me

홈 › Papers

TiFRe: Text-guided Video Frame Reduction for Efficient Video Multi-modal Large Language Models

2026-02-09 · Xiangtian Zheng, Zishuo Wang, Yuxin Peng arxiv

With the rapid development of Large Language Models (LLMs), Video Multi-Modal Large Language Models (Video MLLMs) have achieved remarkable performance in video-language tasks such as video understanding and question answering. However, Video MLLMs face high computational costs, particularly in processing numerous video frames as input, which leads to significant attention computation overhead. A straightforward approach to reduce computational costs is to decrease the number of input video frames. However, simply selecting key frames at a fixed frame rate (FPS) often overlooks valuable information in non-key frames, resulting in notable performance degradation. To address this, we propose Text-guided Video Frame Reduction (TiFRe), a framework that reduces input frames while preserving essential video information. TiFRe uses a Text-guided Frame Sampling (TFS) strategy to select key frames based on user input, which is processed by an LLM to generate a CLIP-style prompt. Pre-trained CLIP encoders calculate the semantic similarity between the prompt and each frame, selecting the most relevant frames as key frames. To preserve video semantics, TiFRe employs a Frame Matching and Merging (FMM) mechanism, which integrates non-key frame information into the selected key frames, minimizing information loss. Experiments show that TiFRe effectively reduces computational costs while improving performance on video-language tasks.

📄 PDF Abstract BibTeX arXiv:2602.08861

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic SimilarityQuestion Answering

Similar Papers 제목 키워드 기반

MotifRetro: Exploring the Combinability-Consistency Trade-offs in retrosynthesis via Dynamic Motif Editing

2023-05-20 · Zhangyang Gao, Xingran Chen, Cheng Tan, Stan Z. Li

Is there a unified framework for graph-based retrosynthesis prediction? Through analysis of full-, semi-, and non-template retrosynthesis methods, we discovered that they strive to strike an optimal balance between combi…

Retrosynthesis

ArtiFree: Detecting and Reducing Generative Artifacts in Diffusion-based Speech Enhancement

2025-09-23 · Bhawana Chhaglani, Yang Gao, Julius Richter, Xilin Li 외 arxiv

Diffusion-based speech enhancement (SE) achieves natural-sounding speech and strong generalization, yet suffers from key limitations like generative artifacts and high inference latency. In this work, we systematically s…

Speech Enhancement

Exact Decomposition of Multifrequency Discrete Real and Complex Signals

2022-02-08 · BaoGuo Liu

'The spectral leakage from windowing and the picket fence effect from discretization' have been among the standard contents in textbooks for many decades. The spectral leakage and picket fence effect would cause the dist…

Deep learning enhanced Rydberg multifrequency microwave recognition

2022-02-28 · Zong-Kai Liu, Li-Hua Zhang, Bang Liu, Zheng-Yuan Zhang 외

Recognition of multifrequency microwave (MW) electric fields is challenging because of the complex interference of multifrequency fields in practical applications. Rydberg atom-based measurements for multifrequency MW el…

Deep Learning

Fourier Analysis on Transient Imaging with a Multifrequency Time-of-Flight Camera

2014-06-01 · CVPR 2014 6 · Jingyu Lin, Yebin Liu, Matthias B. Hullin, Qionghai Dai

A transient image is the optical impulse response of a scene which visualizes light propagation during an ultra-short time interval. In this paper we discover that the data captured by a multifrequency time-of-flight (To…