paper-with-me

Papers

Generalization in Multimodal Language Learning from Simulation

2021-08-03 · Aaron Eisermann, Jae Hee Lee, Cornelius Weber, Stefan Wermter

Neural networks can be powerful function approximators, which are able to model high-dimensional feature distributions from a subset of examples drawn from the target distribution. Naturally, they perform well at generalizing within the limits of their target function, but they often fail to generalize outside of the explicitly learned feature space. It is therefore an open research topic whether and how neural network-based architectures can be deployed for systematic reasoning. Many studies have shown evidence for poor generalization, but they often work with abstract data or are limited to single-channel input. Humans, however, learn and interact through a combination of multiple sensory modalities, and rarely rely on just one. To investigate compositional generalization in a multimodal setting, we generate an extensible dataset with multimodal input sequences from simulation. We investigate the influence of the underlying training data distribution on compostional generalization in a minimal LSTM-based network trained in a supervised, time continuous setting. We find compositional generalization to fail in simple setups while improving with the number of objects, actions, and particularly with a lot of color overlaps between objects. Furthermore, multimodality strongly improves compositional generalization in settings where a pure vision model struggles to generalize.

📄 PDF Abstract BibTeX arXiv:2108.02319

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Realistic Corner Case Generation for Autonomous Vehicles with Multimodal Large Language Model

2024-11-29 · QIUJING LU, Meng Ma, Ximiao Dai, Xuanhan Wang 외

To guarantee the safety and reliability of autonomous vehicle (AV) systems, corner cases play a crucial role in exploring the system's behavior under rare and challenging conditions within simulation environments. Howeve…

Autonomous VehiclesLanguage ModelingLanguage ModellingLarge Language Model+2

VIMA: General Robot Manipulation with Multimodal Prompts

2022-10-06 · Yunfan Jiang, Agrim Gupta, Zichen Zhang, Guanzhi Wang 외

Prompt-based learning has emerged as a successful paradigm in natural language processing, where a single general-purpose language model can be instructed to perform any task specified by input prompts. Yet task specific…

Imitation LearningLanguage ModellingRobot ManipulationSystematic Generalization+1

URDF-Anything: Constructing Articulated Objects with 3D Multimodal Language Model

2025-11-02 · Zhe Li, Xiang Bai, Jieyu Zhang, Zhuangzhe Wu 외 arxiv

Constructing accurate digital twins of articulated objects is essential for robotic simulation training and embodied AI world model building, yet historically requires painstaking manual modeling or multi-stage pipelines…

Parameter Prediction

Frequency-Aware Vision-Language Multimodality Generalization Network for Remote Sensing Image Classification

2025-11-13 · Junjie Zhang, Feng Zhao, Hanqiang Liu, Jun Yu arxiv

The booming remote sensing (RS) technology is giving rise to a novel multimodality generalization task, which requires the model to overcome data heterogeneity while possessing powerful cross-scene generalization ability…

Remote Sensing Image Classification

On the generalization capacity of neural networks during generic multimodal reasoning

2024-01-26 · Takuya Ito, Soham Dan, Mattia Rigotti, James Kozloski 외

The advent of the Transformer has led to the development of large language models (LLM), which appear to demonstrate human-like capabilities. To assess the generality of this class of models and a variety of other base n…

Multimodal ReasoningSystematic Generalization