paper-with-me

홈 › Papers

pNLP-Mixer: an Efficient all-MLP Architecture for Language

2022-02-09 · Francesco Fusco, Damian Pascual, Peter Staar, Diego Antognini

Large pre-trained language models based on transformer architecture have drastically changed the natural language processing (NLP) landscape. However, deploying those models for on-device applications in constrained devices such as smart watches is completely impractical due to their size and inference cost. As an alternative to transformer-based architectures, recent work on efficient NLP has shown that weight-efficient models can attain competitive performance for simple tasks, such as slot filling and intent classification, with model sizes in the order of the megabyte. This work introduces the pNLP-Mixer architecture, an embedding-free MLP-Mixer model for on-device NLP that achieves high weight-efficiency thanks to a novel projection layer. We evaluate a pNLP-Mixer model of only one megabyte in size on two multi-lingual semantic parsing datasets, MTOP and multiATIS. Our quantized model achieves 99.4% and 97.8% the performance of mBERT on MTOP and multi-ATIS, while using 170x fewer parameters. Our model consistently beats the state-of-the-art of tiny models (pQRNN), which is twice as large, by a margin up to 7.8% on MTOP.

📄 PDF Abstract BibTeX arXiv:2202.04350

Code (1)

mindslab-ai/pnlp-mixer pytorch

Tasks

Allintent-classificationIntent ClassificationSemantic Parsingslot-fillingSlot Filling

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
mBERT mBERT
Adam 설명 없음
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…
Linear Warmup With Linear Decay Linear Warmup With Linear Decay is a learning rate schedule in which we increase the learning rate linearly for $n$ updates and then linearly decay afterwards.
WordPiece 설명 없음

Similar Papers 제목 키워드 기반

DeepNLPF: A Framework for Integrating Third Party NLP Tools

2020-05-01 · LREC 2020 5 · Francisco Rodrigues, Rinaldo Lima, William Domingues, Robson Fidalgo 외

Natural Language Processing (NLP) of textual data is usually broken down into a sequence of several subtasks, where the output of one the subtasks becomes the input to the following one, which constitutes an NLP pipeline…

Management

Implementing a Portable Clinical NLP System with a Common Data Model - a Lisp Perspective

2018-11-15 · Yuan Luo, Peter Szolovits

This paper presents a Lisp architecture for a portable NLP system, termed LAPNLP, for processing clinical notes. LAPNLP integrates multiple standard, customized and in-house developed NLP tools. Our system facilitates po…

Computational PhenotypingDomain AdaptationRelation Extraction

CogCompNLP: Your Swiss Army Knife for NLP

2018-05-01 · LREC 2018 5 · Daniel Khashabi, Mark Sammons, Ben Zhou, Tom Redman 외
Semantic Role Labeling

A Unified Understanding of Deep NLP Models for Text Classification

2022-06-19 · Zhen Li, Xiting Wang, Weikai Yang, Jing Wu 외

The rapid development of deep natural language processing (NLP) models for text classification has led to an urgent need for a unified understanding of these models proposed individually. Existing methods cannot meet the…

Classificationtext-classificationText Classification

CICBUAPnlp: Graph-Based Approach for Answer Selection in Community Question Answering Task

2015-06-01 · SEMEVAL 2015 6 · Helena Gomez, Darnes Vilari{\~n}o, David Pinto, Grigori Sidorov
Answer SelectionCommunity Question AnsweringLearning-To-RankNatural Language Inference+2