paper-with-me

Papers

Exploring Meaning Encoded in Random Character Sequences with Character-Aware Language Models

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Natural language processing models learn word representations based on the distributional hypothesis, which asserts that word context (e.g., co-occurrence) correlates with semantic meaning. We propose that n-grams composed of random character sequences, or garble, provide a novel context for studying word meaning both within and beyond extant language. In particular, randomly-generated character n-grams lack semantic meaning but contain primitive information based on the distribution of characters they contain. By studying the embeddings of a large corpus of garble, extant language, and pseudowords using CharacterBERT, we identify an axis in the model's high-dimensional embedding space that separates these classes of n-grams. Furthermore, we show that this axis relates to structure within extant language, including word part of speech, morphology, and concreteness. Thus, in contrast to studies that are mainly limited to extant language, our work reveals that semantic meaning and primitive information are intrinsically linked.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

CharacterBERT 설명 없음

Similar Papers 제목 키워드 기반

Signal in Noise: Exploring Meaning Encoded in Random Character Sequences with Character-Aware Language Models

2022-03-15 · ACL 2022 5 · Mark Chu, Bhargav Srinivasa Desikan, Ethan O. Nadler, D. Ruggiero Lo Sardo 외

Natural language processing models learn word representations based on the distributional hypothesis, which asserts that word context (e.g., co-occurrence) correlates with meaning. We propose that $n$-grams composed of r…

Enabling Cognitive Intelligence Queries in Relational Databases using Low-dimensional Word Embeddings

2016-03-23 · Rajesh Bordawekar, Oded Shmueli

We apply distributed language embedding methods from Natural Language Processing to assign a vector to each database entity associated token (for example, a token may be a word occurring in a table row, or the name of a …

Word Embeddings

A Causal World Model Underlying Next Token Prediction: Exploring GPT in a Controlled Environment

2024-12-10 · Raanan Y. Rohekar, Yaniv Gurwicz, Sungduk Yu, Estelle Aflalo 외

Do generative pre-trained transformer (GPT) models, trained only to predict the next token, implicitly learn a world model from which a sequence is generated one token at a time? We address this question by deriving a ca…

model

A Neuromorphic Model of Learning Meaningful Sequences with Long-Term Memory

2025-09-16 · Laxmi R. Iyer, Ali A. Minai arxiv

Learning meaningful sentences is different from learning a random set of words. When humans understand the meaning, the learning occurs relatively quickly. What mechanisms enable this to happen? In this paper, we examine…

Does Character-level Information Always Improve DRS-based Semantic Parsing?

2023-06-04 · Tomoya Kurosawa, Hitomi Yanaka

Even in the era of massive language models, it has been suggested that character-level representations improve the performance of neural models. The state-of-the-art neural semantic parser for Discourse Representation St…

Semantic Parsing