paper-with-me

홈 › Papers

MLP Architectures for Vision-and-Language Modeling: An Empirical Study

2021-12-08 · Yixin Nie, Linjie Li, Zhe Gan, Shuohang Wang, Chenguang Zhu, Michael Zeng, Zicheng Liu, Mohit Bansal, Lijuan Wang

We initiate the first empirical study on the use of MLP architectures for vision-and-language (VL) fusion. Through extensive experiments on 5 VL tasks and 5 robust VQA benchmarks, we find that: (i) Without pre-training, using MLPs for multimodal fusion has a noticeable performance gap compared to transformers; (ii) However, VL pre-training can help close the performance gap; (iii) Instead of heavy multi-head attention, adding tiny one-head attention to MLPs is sufficient to achieve comparable performance to transformers. Moreover, we also find that the performance gap between MLPs and transformers is not widened when being evaluated on the harder robust VQA benchmarks, suggesting using MLPs for VL fusion can generalize roughly to a similar degree as using transformers. These results hint that MLPs can effectively learn to align vision and text features extracted from lower-level encoders without heavy reliance on self-attention. Based on this, we ask an even bolder question: can we have an all-MLP architecture for VL modeling, where both VL fusion and the vision encoder are replaced with MLPs? Our result shows that an all-MLP VL model is sub-optimal compared to state-of-the-art full-featured VL models when both of them get pre-trained. However, pre-training an all-MLP can surprisingly achieve a better average score than full-featured transformer models without pre-training. This indicates the potential of large-scale pre-training of MLP-like architectures for VL modeling and inspires the future research direction on simplifying well-established VL modeling with less inductive design bias. Our code is publicly available at: https://github.com/easonnie/mlp-vil

📄 PDF Abstract BibTeX arXiv:2112.04453

Code (1)

easonnie/mlp-vil 공식 구현

Tasks

Language ModelingLanguage ModellingVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Efficient Post-training Quantization with FP8 Formats

2023-09-26 · Haihao Shen, Naveen Mellempudi, Xin He, Qun Gao 외

Recent advances in deep learning methods such as LLMs and Diffusion models have created a need for improved quantization methods that can meet the computational demands of these modern architectures while maintaining acc…

image-classificationImage ClassificationLanguage ModelingLanguage Modelling+3

On End-to-End Program Generation from User Intention by Deep Neural Networks

2015-10-25 · Lili Mou, Rui Men, Ge Li, Lu Zhang 외

This paper envisions an end-to-end program generation scenario using recurrent neural networks (RNNs): Users can express their intention in natural language; an RNN then automatically generates corresponding code in a ch…

Understanding Architectures Learnt by Cell-based Neural Architecture Search

2019-09-20 · ICLR 2020 1 · Yao Shu, Wei Wang, Shaofeng Cai

Neural architecture search (NAS) searches architectures automatically for given tasks, e.g., image classification and language modeling. Improving the search efficiency and effectiveness have attracted increasing attenti…

image-classificationImage ClassificationLanguage ModelingLanguage Modelling+1

Empirical Evaluation of Knowledge Distillation from Transformers to Subquadratic Language Models

2025-04-19 · Patrick Haller, Jonas Golde, Alan Akbik

Knowledge distillation is a widely used technique for compressing large language models (LLMs) by training a smaller student model to mimic a larger teacher model. Typically, both the teacher and student are Transformer-…

Knowledge DistillationState Space ModelsTransfer Learning

On DeepSeekMoE: Statistical Benefits of Shared Experts and Normalized Sigmoid Gating

2025-05-16 · Huy Nguyen, Thong T. Doan, Quang Pham, Nghi D. Q. Bui 외

Mixture of experts (MoE) methods are a key component in most large language model architectures, including the recent series of DeepSeek models. Compared to other MoE implementations, DeepSeekMoE stands out because of tw…

Language ModelingLanguage ModellingLarge Language ModelMixture-of-Experts