paper-with-me

홈 › Papers

What is needed for simple spatial language capabilities in VQA?

2019-08-17 · Alexander Kuhnle, Ann Copestake

Visual question answering (VQA) comprises a variety of language capabilities. The diagnostic benchmark dataset CLEVR has fueled progress by helping to better assess and distinguish models in basic abilities like counting, comparing and spatial reasoning in vitro. Following this approach, we focus on spatial language capabilities and investigate the question: what are the key ingredients to handle simple visual-spatial relations? We look at the SAN, RelNet, FiLM and MC models and evaluate their learning behavior on diagnostic data which is solely focused on spatial relations. Via comparative analysis and targeted model modification we identify what really is required to substantially improve upon the CNN-LSTM baseline.

📄 PDF Abstract BibTeX arXiv:1908.06336

Code (0)

등록된 구현이 없습니다.

Tasks

DiagnosticQuestion AnsweringSpatial ReasoningVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Common-Description Learning: A Framework for Learning Algorithms and Generating Subproblems from Few Examples

2016-05-01 · Basem G. El-Barashy

Current learning algorithms face many difficulties in learning simple patterns and using them to learn more complex ones. They also require more examples than humans do to learn the same pattern, assuming no prior knowle…

What's left can't be right -- The remaining positional incompetence of contrastive vision-language models

2023-11-20 · Nils Hoehing, Ellen Rushe, Anthony Ventresque

Contrastive vision-language models like CLIP have been found to lack spatial understanding capabilities. In this paper we discuss the possible causes of this phenomenon by analysing both datasets and embedding space. By …

From Spatial Relations to Spatial Configurations

2020-07-19 · LREC 2020 5 · Soham Dan, Parisa Kordjamshidi, Julia Bonn, Archna Bhatia 외

Spatial Reasoning from language is essential for natural language understanding. Supporting it requires a representation scheme that can capture spatial phenomena encountered in language as well as in images and videos. …

Abstract Meaning RepresentationNatural Language UnderstandingSpatial Reasoning

Rigor in AI: Doing Rigorous AI Work Requires a Broader, Responsible AI-Informed Conception of Rigor

2025-06-17 · Alexandra Olteanu, Su Lin Blodgett, Agathe Balayn, Angelina Wang 외

In AI research and practice, rigor remains largely understood in terms of methodological rigor -- such as whether mathematical, statistical, or computational methods are correctly applied. We argue that this narrow conce…

Agentic AI in Engineering and Manufacturing: Industry Perspectives on Utility, Adoption, Challenges, and Opportunities

2026-03-19 · Kristen M. Edwards, Maxwell Bauer, Claire Jacquillat, A. John Hart 외 arxiv

This work examines how AI, especially agentic systems, is being adopted in engineering and manufacturing workflows, what value it provides today, and what is needed for broader deployment. This is an exploratory and qual…