paper-with-me

홈 › Papers

Free-MoRef: Instantly Multiplexing Context Perception Capabilities of Video-MLLMs within Single Inference

2025-08-04 · Kuo Wang, Quanlong Zheng, Junlin Xie, Yanhao Zhang, Jinguo Luo, Haonan Lu, Liang Lin, Fan Zhou, Guanbin Li arxiv

Video Multimodal Large Language Models~(Video-MLLM) have achieved remarkable advancements in video understanding tasks. However, constrained by the context length limitation in the underlying LLMs, existing Video-MLLMs typically exhibit suboptimal performance on long video scenarios. To understand extended input frames, common solutions span token compression and streaming inference techniques, which sacrifice feature granularity or inference efficiency. Differently, to efficiently achieve comprehensive understanding of longer frame inputs, we draw ideas from MoE and propose a training-free approach \textbf{Free-MoRef}, which instantly multiplexes the context perception capabilities of Video-MLLMs within one inference pass. Specifically, Free-MoRef reconstructs the vision tokens into several short sequences as multi-references. Subsequently, we introduce MoRef-attention, which gathers clues from the multi-reference chunks in parallel to summarize unified query activations. After the shadow layers in LLMs, a reference fusion step is derived to compose a final mixed reasoning sequence with key tokens from parallel chunks, which compensates the cross-reference vision interactions that are neglected in MoRef-attention. By splitting and fusing the long vision token sequences, Free-MoRef achieves improved performance under much lower computing costs in reasoning multiplexed context length, demonstrating strong efficiency and effectiveness. Experiments on VideoMME, MLVU, LongVideoBench show that Free-MoRef achieves full perception of 2$\times$ to 8$\times$ longer input frames without compression on a single A100 GPU while keeping instant responses, thereby bringing significant performance gains, even surpassing dedicatedly trained long-video-MLLMs. Codes are available at https://github.com/wkfdb/Free-MoRef

📄 PDF Abstract BibTeX arXiv:2508.02134

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

World Models Increase Autonomy in Reinforcement Learning

2024-08-19 · Zhao Yang, Thomas M. Moerland, Mike Preuss, Aske Plaat 외

Reinforcement learning (RL) is an appealing paradigm for training intelligent agents, enabling policy acquisition from the agent's own autonomously acquired experience. However, the training process of RL is far from aut…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

MoReFlow: Motion Retargeting Learning through Unsupervised Flow Matching

2025-09-29 · Wontaek Kim, Tianyu Li, Sehoon Ha arxiv

Motion retargeting holds a premise of offering a larger set of motion data for characters and robots with different morphologies. Many prior works have approached this problem via either handcrafted constraints or paired…

Enhanced Atmospheric Turbulence Resiliency with Successive Interference Cancellation DSP in Mode Division Multiplexing Free-Space Optical Links

2022-07-19 · Yiming Li, Zhaozhong Chen, Zhouyi Hu, David M. Benton 외

We experimentally demonstrate the enhanced atmospheric turbulence resiliency in a 137.8 Gbit/s/mode mode-division multiplexing free-space optical communication link through the application of a successive interference ca…

BDG inequalities and their applications for model-free continuous price paths with instant enforcement

2021-09-16 · Rafał M. Łochowski

Shafer and Vovk introduce in their book \cite{ShaferVovk:2018} the notion of \emph{instant enforcement} and \emph{instantly blockable} properties. However, they do not associate these notions with any outer measure, unli…

The OoO VLIW JIT Compiler for GPU Inference

2019-01-28 · Paras Jain, Xiangxi Mo, Ajay Jain, Alexey Tumanov 외

Current trends in Machine Learning~(ML) inference on hardware accelerated devices (e.g., GPUs, TPUs) point to alarmingly low utilization. As ML inference is increasingly time-bounded by tight latency SLOs, increasing dat…

GPU