paper-with-me

홈 › Papers

VEN-VL: A Visual Ensemble MoE Framework for Effective and Efficient Multi-Modal Understanding

2026-05-25 · Yinghao Wu, Zhuoyan Luo, Yiyao Yu, Zhaojian Yu, Yujiu Yang, Xiao-Ping Zhang arxiv

Despite the remarkable progress achieved by recent efficient methods in accelerating multimodal understanding, they still suffer from noticeable performance degradation. Their emphasis on the high compression ratio of a single visual clue and reliance on the heuristic pruning strategy with coarse attention alignment incurs a bottleneck on the information capacity and density of visual tokens. Addressing this limitation, we propose VEN-VL, a visual ensemble MoE framework for effective and efficient perception following the enrich then compact principle. Specifically, we first enrich the information capacity by unifying the visual representations of different perspectives, and then progressively compact it with adaptive routers in specialized visual experts to enhance the information density. Furthermore, we incorporate the reconstruction ability of vanilla structure via explicit visual supervision, facilitating crucial information preservation. Experimental results demonstrate our superiority in complex visual tasks with few information-condensed tokens, which effectively bridges the gap between performance and efficiency.

📄 PDF Abstract BibTeX arXiv:2605.25952

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CrossGET: Cross-Guided Ensemble of Tokens for Accelerating Vision-Language Transformers

2023-05-27 · Dachuan Shi, Chaofan Tao, Anyi Rao, Zhendong Yang 외

Recent vision-language models have achieved tremendous advances. However, their computational costs are also escalating dramatically, making model acceleration exceedingly critical. To pursue more efficient vision-langua…

Image CaptioningImage RetrievalImage-text RetrievalImage-to-Text Retrieval+5

Hitachi at SemEval-2020 Task 8: Simple but Effective Modality Ensemble for Meme Emotion Recognition

2020-12-01 · SEMEVAL 2020 · Terufumi Morishita, Gaku Morio, Shota Horiguchi, Hiroaki Ozaki 외

Users of social networking services often share their emotions via multi-modal content, usually images paired with text embedded in them. SemEval-2020 task 8, Memotion Analysis, aims at automatically recognizing these em…

Emotion Recognition

CARPE: Context-Aware Image Representation Prioritization via Ensemble for Large Vision-Language Models

2026-01-20 · Donghee Lee, Rui Cai, Zhe Zhao arxiv

Large vision-language models (LVLMs) are typically trained using autoregressive language modeling objectives, which align visual representations with linguistic space. While effective for multimodal reasoning, this align…

Multimodal ReasoningImage Classification

Cross-Lingual Cross-Modal Consolidation for Effective Multilingual Video Corpus Moment Retrieval

2022-07-01 · Findings (NAACL) 2022 7 · Jiaheng Liu, Tan Yu, Hanyu Peng, Mingming Sun 외

Existing multilingual video corpus moment retrieval (mVCMR) methods are mainly based on a two-stream structure. The visual stream utilizes the visual content in the video to estimate the query-visual similarity, and the …

Moment RetrievalRetrievalVideo Corpus Moment RetrievalVideo Similarity

FusionEnsemble-Net: An Attention-Based Ensemble of Spatiotemporal Networks for Multimodal Sign Language Recognition

2025-08-12 · Md. Milon Islam, Md Rezwanul Haque, S M Taslim Uddin Raju, Fakhri Karray arxiv

Accurate recognition of sign language in healthcare communication poses a significant challenge, requiring frameworks that can accurately interpret complex multimodal gestures. To deal with this, we propose FusionEnsembl…

Sign Language RecognitionGesture Recognition