paper-with-me

홈 › Papers

Neologism Learning for Controllability and Self-Verbalization

2025-10-09 · John Hewitt, Oyvind Tafjord, Robert Geirhos, Been Kim arxiv

Humans invent new words when there is a rising demand for a new useful concept (e.g., doomscrolling). We explore and validate a similar idea in our communication with LLMs: introducing new words to better understand and control the models, expanding on the recently introduced neologism learning. This method introduces a new word by adding a new word embedding and training with examples that exhibit the concept with no other changes in model parameters. We show that adding a new word allows for control of concepts such as flattery, incorrect answers, text length, as well as more complex concepts in AxBench. We discover that neologisms can also further our understanding of the model via self-verbalization: models can describe what each new word means to them in natural language, like explaining that a word that represents a concept of incorrect answers means ``a lack of complete, coherent, or meaningful answers...'' To validate self-verbalizations, we introduce plug-in evaluation: we insert the verbalization into the context of a model and measure whether it controls the target concept. In some self-verbalizations, we find machine-only synonyms: words that seem unrelated to humans but cause similar behavior in machines. Finally, we show how neologism learning can jointly learn multiple concepts in multiple words.

📄 PDF Abstract BibTeX arXiv:2510.08506

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Neologism Learning as a Parameter-Efficient Alternative to Fine-Tuning for Model Steering

2025-12-21 · Sungjoon Park, Varun Ramamurthi, Owen Terry arxiv

In language modeling, neologisms are new tokens trained to represent a concept not already included in a given model's vocabulary. Neologisms can be used to encourage specific behavior in models, for example by appending…

CNeo-Bench: Diagnosing Large Language Models on Chinese Neologisms

2026-08-28 · Kaiyan Zhao, Zhongtao Miao, Zheyong Xie, Shaosheng Cao 외 arxiv

Chinese neologisms exploit diverse and unique linguistic mechanisms, such as phonetic substitution (e.g., 886 for ``bye-bye'') and visual character decomposition that are rare in other languages. We introduce CNeo-Bench,…

Classification and Analysis of Neologisms Produced by Learners of Spanish: Effects of Proficiency and Task

2020-07-01 · WS 2020 7 · Shira Wein

The Spanish Learner Language Oral Corpora (SPLLOC) of transcribed conversations between investigators and language learners contains a set of neologism tags. In this work, the utterances tagged as neologisms are broken d…

General ClassificationLanguage Acquisition

NEO-BENCH: Evaluating Robustness of Large Language Models with Neologisms

2024-02-19 · Jonathan Zheng, Alan Ritter, Wei Xu

The performance of Large Language Models (LLMs) degrades from the temporal drift between data used for model training and newer text seen during inference. One understudied avenue of language change causing data drift is…

Machine TranslationNatural Language UnderstandingSentence

From 124 Million Tokens to 1,021 Neologisms: A Large-Scale Pipeline for Automatic Neologism Detection

2026-05-07 · Diego Rossini, Lonneke van der Plas arxiv

We present a scalable, modular pipeline for automatic neologism detection that combines rule-based filtering with LLM classification. The pipeline is grounded in two complementary word-formation frameworks, grammatical a…