paper-with-me

Papers

ViLPAct: A Benchmark for Compositional Generalization on Multimodal Human Activities

2022-10-11 · Terry Yue Zhuo, Yaqing Liao, Yuecheng Lei, Lizhen Qu, Gerard de Melo, Xiaojun Chang, Yazhou Ren, Zenglin Xu

We introduce ViLPAct, a novel vision-language benchmark for human activity planning. It is designed for a task where embodied AI agents can reason and forecast future actions of humans based on video clips about their initial activities and intents in text. The dataset consists of 2.9k videos from \charades extended with intents via crowdsourcing, a multi-choice question test set, and four strong baselines. One of the baselines implements a neurosymbolic approach based on a multi-modal knowledge base (MKB), while the other ones are deep generative models adapted from recent state-of-the-art (SOTA) methods. According to our extensive experiments, the key challenges are compositional generalization and effective use of information from both modalities.

📄 PDF Abstract BibTeX arXiv:2210.05556

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Test 설명 없음
BASE 설명 없음

Similar Papers 제목 키워드 기반

On the generalization capacity of neural networks during generic multimodal reasoning

2024-01-26 · Takuya Ito, Soham Dan, Mattia Rigotti, James Kozloski 외

The advent of the Transformer has led to the development of large language models (LLM), which appear to demonstrate human-like capabilities. To assess the generality of this class of models and a variety of other base n…

Multimodal ReasoningSystematic Generalization

CrossMed: A Multimodal Cross-Task Benchmark for Compositional Generalization in Medical Imaging

2025-11-14 · Pooja Singh, Siddhant Ujjain, Tapan Kumar Gandhi, Sandeep Kumar arxiv

Recent advances in multimodal large language models have enabled unified processing of visual and textual inputs, offering promising applications in general-purpose medical AI. However, their ability to generalize compos…

Visual Question Answering

Generalization in Multimodal Language Learning from Simulation

2021-08-03 · Aaron Eisermann, Jae Hee Lee, Cornelius Weber, Stefan Wermter

Neural networks can be powerful function approximators, which are able to model high-dimensional feature distributions from a subset of examples drawn from the target distribution. Naturally, they perform well at general…

Syntax-Guided Transformers: Elevating Compositional Generalization and Grounding in Multimodal Environments

2023-11-07 · Danial Kamali, Parisa Kordjamshidi

Compositional generalization, the ability of intelligent models to extrapolate understanding of components to novel compositions, is a fundamental yet challenging facet in AI research, especially within multimodal enviro…

Compositional Generalization (AVG)Dependency Parsing

COLA: A Benchmark for Compositional Text-to-image Retrieval

2023-05-05 · NeurIPS 2023 11 · Arijit Ray, Filip Radenovic, Abhimanyu Dubey, Bryan A. Plummer 외

Compositional reasoning is a hallmark of human visual intelligence. Yet, despite the size of large vision-language models, they struggle to represent simple compositions by combining objects with their attributes. To mea…

AttributeCoLAImage RetrievalRetrieval