paper-with-me

홈 › Papers

How Powerful are Decoder-Only Transformer Neural Models?

2023-05-26 · Jesse Roberts

In this article we prove that the general transformer neural model undergirding modern large language models (LLMs) is Turing complete under reasonable assumptions. This is the first work to directly address the Turing completeness of the underlying technology employed in GPT-x as past work has focused on the more expressive, full auto-encoder transformer architecture. From this theoretical analysis, we show that the sparsity/compressibility of the word embedding is an important consideration for Turing completeness to hold. We also show that Transformers are are a variant of B machines studied by Hao Wang.

📄 PDF Abstract BibTeX arXiv:2305.17026

Code (0)

등록된 구현이 없습니다.

Tasks

Decoder

Similar Papers 제목 키워드 기반

When are 1.58 bits enough? A Bottom-up Exploration of BitNet Quantization

2024-11-08 · Jacob Nielsen, Lukas Galke, Peter Schneider-Kamp

Contemporary machine learning models, such as language models, are powerful, but come with immense resource requirements both at training and inference time. It has been shown that decoder-only language models can be tra…

DecoderQuantization

Efficient Encoder-Decoder Transformer Decoding for Decomposable Tasks

2024-03-19 · Bo-Ru Lu, Nikita Haduong, Chien-Yu Lin, Hao Cheng 외

Transformer-based NLP models are powerful but have high computational costs that limit deployment. Finetuned encoder-decoder models are popular in specialized domains and can outperform larger more generalized decoder-on…

DecoderDialogue State TrackingQuestion Answering

iSegFormer: Interactive Segmentation via Transformers with Application to 3D Knee MR Images

2021-12-21 · Qin Liu, Zhenlin Xu, Yining Jiao, Marc Niethammer

We propose iSegFormer, a memory-efficient transformer that combines a Swin transformer with a lightweight multilayer perceptron (MLP) decoder. With the efficient Swin transformer blocks for hierarchical self-attention an…

DecoderImage SegmentationInteractive SegmentationMedical Image Segmentation+1

Instantaneous Grammatical Error Correction with Shallow Aggressive Decoding

2021-06-09 · ACL 2021 5 · Xin Sun, Tao Ge, Furu Wei, Houfeng Wang

In this paper, we propose Shallow Aggressive Decoding (SAD) to improve the online inference efficiency of the Transformer for instantaneous Grammatical Error Correction (GEC). SAD optimizes the online inference efficienc…

DecoderGrammatical Error Correction

Efficiently Summarizing Text and Graph Encodings of Multi-Document Clusters

2021-06-01 · NAACL 2021 4 · Ramakanth Pasunuru, Mengwen Liu, Mohit Bansal, Sujith Ravi 외

This paper presents an efficient graph-enhanced approach to multi-document summarization (MDS) with an encoder-decoder Transformer model. This model is based on recent advances in pre-training both encoder and decoder on…

DecoderDocument SummarizationMulti-Document Summarization