paper-with-me

홈 › Papers

Modulating Bottom-Up and Top-Down Visual Processing via Language-Conditional Filters

2020-03-28 · İlker Kesen, Ozan Arkan Can, Erkut Erdem, Aykut Erdem, Deniz Yuret

How to best integrate linguistic and perceptual processing in multi-modal tasks that involve language and vision is an important open problem. In this work, we argue that the common practice of using language in a top-down manner, to direct visual attention over high-level visual features, may not be optimal. We hypothesize that the use of language to also condition the bottom-up processing from pixels to high-level features can provide benefits to the overall performance. To support our claim, we propose a U-Net-based model and perform experiments on two language-vision dense-prediction tasks: referring expression segmentation and language-guided image colorization. We compare results where either one or both of the top-down and bottom-up visual branches are conditioned on language. Our experiments reveal that using language to control the filters for bottom-up visual processing in addition to top-down attention leads to better results on both tasks and achieves competitive performance. Our linguistic analysis suggests that bottom-up conditioning improves segmentation of objects especially when input text refers to low-level visual concepts. Code is available at https://github.com/ilkerkesen/bvpr.

📄 PDF Abstract BibTeX arXiv:2003.12739

Code (1)

ilkerkesen/bvpr 공식 구현 pytorch

Tasks

ColorizationImage ColorizationReferring ExpressionReferring Expression SegmentationSemantic Segmentation

Similar Papers 제목 키워드 기반

Language Controls More Than Top-Down Attention: Modulating Bottom-Up Visual Processing with Referring Expressions

2021-01-01 · Ozan Arkan Can, Ilker Kesen, Deniz Yuret

How to best integrate linguistic and perceptual processing in multimodal tasks is an important open problem. In this work we argue that the common technique of using language to direct visual attention over high-level vi…

Referring Expression

Bootstrapping Top-down Information for Self-modulating Slot Attention

2024-11-04 · Dongwon Kim, Seoyeon Kim, Suha Kwak

Object-centric learning (OCL) aims to learn representations of individual objects within visual scenes without manual supervision, facilitating efficient and effective visual reasoning. Traditional OCL methods primarily …

ObjectObject DiscoveryVisual Reasoning

Modeling Bottom-up Information Quality during Language Processing

2025-09-21 · Cui Ding, Yanning Yin, Lena A. Jäger, Ethan Gotlieb Wilcox arxiv

Contemporary theories model language processing as integrating both top-down expectations and bottom-up inputs. One major prediction of such models is that the quality of the bottom-up inputs modulates ease of processing…

Modulating early visual processing by language

2017-07-02 · NeurIPS 2017 12 · Harm de Vries, Florian Strub, Jérémie Mary, Hugo Larochelle 외

It is commonly assumed that language refers to high-level visual concepts while leaving low-level visual processing unaffected. This view dominates the current literature in computational models for language-vision tasks…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

BUS:Efficient and Effective Vision-language Pre-training with Bottom-Up Patch Summarization

2023-07-17 · Chaoya Jiang, Haiyang Xu, Wei Ye, Qinghao Ye 외

Vision Transformer (ViT) based Vision-Language Pre-training (VLP) models have demonstrated impressive performance in various tasks. However, the lengthy visual token sequences fed into ViT can lead to training inefficien…

DecoderText Summarization