paper-with-me

Papers

Words That Make Language Models Perceive

2025-10-02 · Sophie L. Wang, Phillip Isola, Brian Cheung arxiv

Large language models (LLMs) trained purely on text ostensibly lack any direct perceptual experience, yet their internal representations are implicitly shaped by multimodal regularities encoded in language. We test the hypothesis that explicit sensory prompting can surface this latent structure, bringing a text-only LLM into closer representational alignment with specialist vision and audio encoders. When a sensory prompt tells the model to 'see' or 'hear', it cues the model to resolve its next-token predictions as if they were conditioned on latent visual or auditory evidence that is never actually supplied. Our findings reveal that lightweight prompt engineering can reliably activate modality-appropriate representations in purely text-trained LLMs.

📄 PDF Abstract BibTeX arXiv:2510.02425

Code (0)

등록된 구현이 없습니다.

Tasks

Prompt Engineering

Similar Papers 제목 키워드 기반

Does BERT Learn as Humans Perceive? Understanding Linguistic Styles through Lexica

2021-09-06 · EMNLP 2021 11 · Shirley Anugrah Hayati, Dongyeop Kang, Lyle Ungar

People convey their intention and attitude through linguistic styles of the text that they write. In this study, we investigate lexicon usages across styles throughout two lenses: human perception and machine word import…

Benchmarking

An exploratory study of L1-specific non-words

2020-09-02 · David Alfter

In this paper, we explore L1-specific non-words, i.e. non-words in a target language (in this case Swedish) that are re-ranked by a different-language language model. We surmise that speakers of a certain L1 will react d…

Language ModelingLanguage ModellingRe-Ranking

Modeling Language Vagueness in Privacy Policies using Deep Neural Networks

2018-05-25 · Fei Liu, Nicole Lee Fella, Kexin Liao

Website privacy policies are too long to read and difficult to understand. The over-sophisticated language makes privacy notices to be less effective than they should be. People become even less willing to share their pe…

Unsupervised Language agnostic WER Standardization

2023-03-09 · Satarupa Guha, Rahul Ambavat, Ankur Gupta, Manish Gupta 외

Word error rate (WER) is a standard metric for the evaluation of Automated Speech Recognition (ASR) systems. However, WER fails to provide a fair evaluation of human perceived quality in presence of spelling variations, …

speech-recognitionSpeech RecognitionTransliteration

VideoPerceiver: Enhancing Fine-Grained Temporal Perception in Video Multimodal Large Language Models

2025-11-24 · Fufangchen Zhao, Liao Zhang, Daiqi Shi, Yuanjun Gao 외 arxiv

We propose VideoPerceiver, a novel video multimodal large language model (VMLLM) that enhances fine-grained perception in video understanding, addressing VMLLMs' limited ability to reason about brief actions in short cli…

Reinforcement LearningAction Understanding