paper-with-me

Papers

Character Sequence-to-Sequence Model with Global Attention for Universal Morphological Reinflection

2017-08-01 · CONLL 2017 8 · Qile Zhu, Yanjun Li, Xiaolin Li
📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationQuestion Answering

Similar Papers 제목 키워드 기반

Universal Approximation with Softmax Attention

2025-04-22 · Jerry Yao-Chieh Hu, Hude Liu, Hong-Yu Chen, Weimin Wu 외

We prove that with linear transformations, both (i) two-layer self-attention and (ii) one-layer self-attention followed by a softmax function are universal approximators for continuous sequence-to-sequence functions on c…

Big Bird: Transformers for Longer Sequences

2020-07-28 · NeurIPS 2020 12 · Manzil Zaheer, Guru Guruganesh, Avinava Dubey, Joshua Ainslie 외

Transformers-based models, such as BERT, have been one of the most successful deep learning models for NLP. Unfortunately, one of their core limitations is the quadratic dependency (mainly in terms of memory) on the sequ…

Linguistic AcceptabilityNatural Language InferenceQuestion AnsweringSemantic Textual Similarity+2

Are Transformers universal approximators of sequence-to-sequence functions?

2019-12-20 · ICLR 2020 1 · Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J. Reddi 외

Despite the widespread adoption of Transformer models for NLP tasks, the expressive power of these models is not well-understood. In this paper, we establish that Transformer models are universal approximators of continu…

Sequence-To-Sequence Domain Adaptation Network for Robust Text Image Recognition

2019-06-01 · CVPR 2019 6 · Yaping Zhang, Shuai Nie, Wenju Liu, Xing Xu 외

Domain adaptation has shown promising advances for alleviating domain shift problem. However, recent visual domain adaptation works usually focus on non-sequential object recognition with a global coarse alignment, which…

DecoderDomain AdaptationObject Recognition

Prompting a Pretrained Transformer Can Be a Universal Approximator

2024-02-22 · Aleksandar Petrov, Philip H. S. Torr, Adel Bibi

Despite the widespread adoption of prompting, prompt tuning and prefix-tuning of transformer models, our theoretical understanding of these fine-tuning methods remains limited. A key question is whether one can arbitrari…