paper-with-me

홈 › Papers

mLongT5: A Multilingual and Efficient Text-To-Text Transformer for Longer Sequences

2023-05-18 · David Uthus, Santiago Ontañón, Joshua Ainslie, Mandy Guo

We present our work on developing a multilingual, efficient text-to-text transformer that is suitable for handling long inputs. This model, called mLongT5, builds upon the architecture of LongT5, while leveraging the multilingual datasets used for pretraining mT5 and the pretraining tasks of UL2. We evaluate this model on a variety of multilingual summarization and question-answering tasks, and the results show stronger performance for mLongT5 when compared to existing multilingual models such as mBART or M-BERT.

📄 PDF Abstract BibTeX arXiv:2305.11129

Code (1)

google-research/longt5 공식 구현 tf

Tasks

Question Answering

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Inverse Square Root Schedule Inverse Square Root is a learning rate schedule 1 / $\sqrt{\max\left(n, k\right)}$ where $n$ is the current training iteration and $k$ is the number of warm-up steps. This…
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
SentencePiece 설명 없음

Similar Papers 제목 키워드 기반

Transformer based Multilingual document Embedding model

2020-08-19 · Wei Li, Brian Mak

One of the current state-of-the-art multilingual document embedding model LASER is based on the bidirectional LSTM neural machine translation model. This paper presents a transformer-based sentence/document embedding mod…

Document EmbeddingMachine TranslationmodelNMT+2

Multilingual and Continuous Backchannel Prediction: A Cross-lingual Study

2025-12-16 · Koji Inoue, Mikey Elmers, Yahui Fu, Zi Haur Pang 외 arxiv

We present a multilingual, continuous backchannel prediction model for Japanese, English, and Chinese, and use it to investigate cross-linguistic timing behavior. The model is Transformer-based and operates at the frame …

Functional Interpolation for Relative Positions Improves Long Context Transformers

2023-10-06 · Shanda Li, Chong You, Guru Guruganesh, Joshua Ainslie 외

Preventing the performance decay of Transformers on inputs longer than those used for training has been an important challenge in extending the context length of these models. Though the Transformer architecture has fund…

Language ModelingLanguage ModellingPosition

Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context

2019-01-09 · ACL 2019 7 · Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell 외

Transformers have a potential of learning longer-term dependency, but are limited by a fixed-length context in the setting of language modeling. We propose a novel neural architecture Transformer-XL that enables learning…

ArticlesLanguage ModelingLanguage Modelling

Augmented Transformers with Adaptive n-grams Embedding for Multilingual Scene Text Recognition

2023-02-28 · Xueming Yan, Zhihang Fang, Yaochu Jin

While vision transformers have been highly successful in improving the performance in image-based tasks, not much work has been reported on applying transformers to multilingual scene text recognition due to the complexi…

Language IdentificationScene Text Recognition