paper-with-me

홈 › Papers

Are Transformers a Modern Version of ELIZA? Observations on French Object Verb Agreement

2021-09-21 · EMNLP 2021 11 · Bingzhi Li, Guillaume Wisniewski, Benoit Crabbé

Many recent works have demonstrated that unsupervised sentence representations of neural networks encode syntactic information by observing that neural language models are able to predict the agreement between a verb and its subject. We take a critical look at this line of research by showing that it is possible to achieve high accuracy on this agreement task with simple surface heuristics, indicating a possible flaw in our assessment of neural networks' syntactic ability. Our fine-grained analyses of results on the long-range French object-verb agreement show that contrary to LSTMs, Transformers are able to capture a non-trivial amount of grammatical structure.

📄 PDF Abstract BibTeX arXiv:2109.10133

Code (0)

등록된 구현이 없습니다.

Tasks

Sentence

Similar Papers 제목 키워드 기반

Were RNNs All We Needed?

2024-10-02 · Leo Feng, Frederick Tung, Mohamed Osama Ahmed, Yoshua Bengio 외

The introduction of Transformers in 2017 reshaped the landscape of deep learning. Originally proposed for sequence modelling, Transformers have since achieved widespread success across various domains. However, the scala…

AllMamba

The Parallelism Tradeoff: Limitations of Log-Precision Transformers

2022-07-02 · William Merrill, Ashish Sabharwal

Despite their omnipresence in modern NLP, characterizing the computational power of transformer neural nets remains an interesting open question. We prove that transformers whose arithmetic precision is logarithmic in th…

Open-Ended Question Answering

FlashRNN: Optimizing Traditional RNNs on Modern Hardware

2024-12-10 · Korbinian Pöppel, Maximilian Beck, Sepp Hochreiter

While Transformers and other sequence-parallelizable neural network architectures seem like the current state of the art in sequence modeling, they specifically lack state-tracking capabilities. These are important for t…

GPULogical Reasoning

Exact Expressive Power of Transformers with Padding

2025-05-25 · William Merrill, Ashish Sabharwal

Chain of thought is a natural inference-time method for increasing the computational power of transformer-based large language models (LLMs), but comes at the cost of sequential decoding. Are there more efficient alterna…

Hard Attention

Detection of Text Reuse in French Medical Corpora

2016-12-01 · WS 2016 12 · Eva D{'}hondt, Cyril Grouin, Aur{\'e}lie N{\'e}v{\'e}ol, Efstathios Stamatatos 외

Electronic Health Records (EHRs) are increasingly available in modern health care institutions either through the direct creation of electronic documents in hospitals{'} health information systems, or through the digitiz…

De-identificationOptical Character Recognition (OCR)