paper-with-me

홈 › Papers

Language Controls More Than Top-Down Attention: Modulating Bottom-Up Visual Processing with Referring Expressions

2021-01-01 · Ozan Arkan Can, Ilker Kesen, Deniz Yuret

How to best integrate linguistic and perceptual processing in multimodal tasks is an important open problem. In this work we argue that the common technique of using language to direct visual attention over high-level visual features may not be optimal. Using language throughout the bottom-up visual pathway, going from pixels to high-level features, may be necessary. Our experiments on several English referring expression datasets show significant improvements when language is used to control the filters for bottom-up visual processing in addition to top-down attention.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Referring Expression

Similar Papers 제목 키워드 기반

Component Ablation for Efficient Hybrid Language Model Architectures: Performance, Resilience, and Compression Implications

2026-03-23 · Hector Borobia, Elies Seguí-Mas, Guillermina Tormo-Carbó arxiv

Hybrid language models combine softmax attention with linear-time sequence mechanisms such as state-space or linear-attention layers, but the functional contribution of each component type remains insufficiently characte…

Representation as a Bottleneck for Mechanistic Interpretability: The Manifestation Unit Protocol

2026-06-30 · Hussein Chouman, Wataru Sasaki, Tomokazu Matsui, Hirohiko Suwa 외 arxiv

Mechanistic interpretability has produced a rich inventory of component-level analyses that characterise what neural-network components encode and how they interact. Their outputs, however, are not easily reusable: selec…

The Attentional White Bear Effect in Transformer Language Models

2026-05-27 · Rebecca Ramnauth, Brian Scassellati arxiv

Instruction-based suppression is widely used to prevent language models from generating prohibited content, yet it remains unclear whether suppression reduces internal representation or merely suppresses expression. We i…

LILA: Language-Informed Latent Actions

2021-11-05 · Siddharth Karamcheti, Megha Srivastava, Percy Liang, Dorsa Sadigh

We introduce Language-Informed Latent Actions (LILA), a framework for learning natural language interfaces in the context of human-robot collaboration. LILA falls under the shared autonomy paradigm: in addition to provid…

Imitation Learning

Semantic Characteristics of Schizophrenic Speech

2019-04-16 · WS 2019 6 · Kfir Bar, Vered Zilberstein, Ido Ziv, Heli Baram 외

Natural language processing tools are used to automatically detect disturbances in transcribed speech of schizophrenia inpatients who speak Hebrew. We measure topic mutation over time and show that controls maintain more…