paper-with-me

홈 › Papers

HyperVL: An Efficient and Dynamic Multimodal Large Language Model for Edge Devices

2025-12-16 · HyperAI Team, Yuchen Liu, Kaiyang Han, Zhiqiang Xia, Yuhang Dong, Chen Song, Kangyu Tang, Jiaming Xu, Xiushi Feng, WenXuan Yu, Li Peng, Mingyang Wang, Kai Wang, Changpeng Yang, Yang Li, Haoyu Lu, Hao Wang, Bingna Xu, Guangyao Liu, Long Huang, Kaibin Guo, Jinyang Wu, Dan Wu, Hongzhen Wang, Peng Zhou, Shuai Nie, Shande Wang, Runyu Shi, Ying Huang arxiv

Current multimodal large lanauge models possess strong perceptual and reasoning capabilities, however high computational and memory requirements make them difficult to deploy directly on on-device environments. While small-parameter models are progressively endowed with strong general capabilities, standard Vision Transformer (ViT) encoders remain a critical bottleneck, suffering from excessive latency and memory consumption when processing high-resolution inputs.To address these challenges, we introduce HyperVL, an efficient multimodal large language model tailored for on-device inference. HyperVL adopts an image-tiling strategy to cap peak memory usage and incorporates two novel techniques: (1) a Visual Resolution Compressor (VRC) that adaptively predicts optimal encoding resolutions to eliminate redundant computation, and (2) Dual Consistency Learning (DCL), which aligns multi-scale ViT encoders within a unified framework, enabling dynamic switching between visual branches under a shared LLM. Extensive experiments demonstrate that HyperVL achieves state-of-the-art performance among models of comparable size across multiple benchmarks. Furthermore, it significantly significantly reduces latency and power consumption on real mobile devices, demonstrating its practicality for on-device multimodal inference.

📄 PDF Abstract BibTeX arXiv:2512.14052

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

HyperVLA: Efficient Inference in Vision-Language-Action Models via Hypernetworks

2025-10-06 · Zheng Xiong, Kang Li, Zilin Wang, Matthew Jackson 외 arxiv

Built upon language and vision foundation models with strong generalization ability and trained on large-scale robotic data, Vision-Language-Action (VLA) models have recently emerged as a promising approach to learning g…

Zero-shot Generalization

HyperVLP: Enhancing Hierarchical Surgical Video-Language Pre-training in Hyperbolic Space

2026-06-30 · Yaojun Hu, Kun Yuan, Nassir Navab, Haochao Ying 외 arxiv

Surgical vision-language foundation models typically adopt educational materials, such as surgical lecture videos, to transfer surgical knowledge encoded in language into visual representations. These knowledge are multi…

Hybrid-DMKG: A Hybrid Reasoning Framework over Dynamic Multimodal Knowledge Graphs for Multimodal Multihop QA with Knowledge Editing

2025-11-30 · Li Yuan, Qingfei Huang, Bingshan Zhu, Yi Cai 외 arxiv

Multimodal Knowledge Editing (MKE) extends traditional knowledge editing to settings involving both textual and visual modalities. However, existing MKE benchmarks primarily assess final answer correctness while neglecti…

Multimodal ReasoningQuestion Answeringknowledge editingKnowledge Graphs

Dynamic Knowledge Integration for Enhanced Vision-Language Reasoning

2025-01-15 · Julian Perry, Surasakdi Siripong, Thanakorn Phonchai

Large Vision-Language Models (LVLMs) have demonstrated impressive capabilities in multimodal tasks, but their performance is often constrained by the lack of external knowledge integration, limiting their ability to hand…

Question AnsweringVisual Question AnsweringWorld Knowledge

KBE-DME: Dynamic Multimodal Evaluation via Knowledge Enhanced Benchmark Evolution

2025-10-24 · Junzhe Zhang, Huixuan Zhang, Xiaojun Wan arxiv

The rapid progress of multimodal large language models (MLLMs) calls for more reliable evaluation protocols. Existing static benchmarks suffer from the potential risk of data contamination and saturation, leading to infl…