paper-with-me

Papers

FiLM: Visual Reasoning with a General Conditioning Layer

2017-09-22 · Ethan Perez, Florian Strub, Harm de Vries, Vincent Dumoulin, Aaron Courville

We introduce a general-purpose conditioning method for neural networks called FiLM: Feature-wise Linear Modulation. FiLM layers influence neural network computation via a simple, feature-wise affine transformation based on conditioning information. We show that FiLM layers are highly effective for visual reasoning - answering image-related questions which require a multi-step, high-level process - a task which has proven difficult for standard deep learning methods that do not explicitly model reasoning. Specifically, we show on visual reasoning tasks that FiLM layers 1) halve state-of-the-art error for the CLEVR benchmark, 2) modulate features in a coherent manner, 3) are robust to ablations and architectural modifications, and 4) generalize well to challenging, new data from few examples or even zero-shot.

📄 PDF Abstract BibTeX arXiv:1709.07871

Code (7)

ethanjperez/film 공식 구현 pytorch
CPJKU/audio_conditioned_unet pytorch
GuessWhatGame/clevr tf
caffeinism/film-pytorch pytorch
jjgo/hyperlight pytorch
kdaip/stabletts pytorch
keonlee9420/Daft-Exprt pytorch

Tasks

Image Retrieval with Multi-Modal QueryVisual Question Answering (VQA)Visual Question Answering (VQA) Split AVisual Question Answering (VQA) Split BVisual Reasoning

Similar Papers 제목 키워드 기반

Visual Reasoning with Multi-hop Feature Modulation

2018-08-03 · ECCV 2018 9 · Florian Strub, Mathieu Seurin, Ethan Perez, Harm de Vries 외

Recent breakthroughs in computer vision and natural language processing have spurred interest in challenging multi-modal tasks such as visual question-answering and visual dialogue. For such tasks, one successful approac…

Question AnsweringVisual DialogVisual Question AnsweringVisual Question Answering (VQA)+1

FiLM-Nav: Efficient and Generalizable Navigation via VLM Fine-tuning

2025-09-19 · Naoki Yokoyama, Sehoon Ha arxiv

Enabling robotic assistants to navigate complex environments and locate objects described in free-form language is a critical capability for real-world deployment. While foundation models, particularly Vision-Language Mo…

Spatial Reasoning

Benefits of Linear Conditioning with Metadata for Image Segmentation

2021-02-18 · Andreanne Lemay, Charley Gros, Olivier Vincent, Yaou Liu 외

Medical images are often accompanied by metadata describing the image (vendor, acquisition parameters) and the patient (disease type or severity, demographics, genomics). This metadata is usually disregarded by image seg…

Image SegmentationMissing LabelsSegmentationSemantic Segmentation+1

EmbodiedDiffusion: End-to-End Traversability-Guided Visual Diffusion for Heterogeneous Robot Navigation

2025-12-02 · Iana Zhura, Sausar Karaf, Faryal Batool, Nipun Dhananjaya Weerakkodi Mudalige 외 arxiv

Visual traversability estimation is central to autonomous navigation, yet most approaches either rely on prompt-driven Vision-Language Model (VLM) or decouple traversability from trajectory planning, requiring separate p…

Prompt Engineering

EquiFiLM: Charge-Conditioned Equivariant Force Fields via Feature-wise Linear Modulation

2026-07-06 · Samuel Sahel-Schackis, Ken-ichi Nomura, Aiichiro Nakano, Matthias F. Kling 외 arxiv

Foundation machine learning force fields (MLFFs) such as MACE-MP-0 and UMA cover broad chemical space at near density functional theory (DFT) accuracy. However, they assume equilibrium ground-state physics and do not nat…