paper-with-me

홈 › Papers

Neural-Symbolic VQA: Disentangling Reasoning from Vision and Language Understanding

2018-10-04 · NeurIPS 2018 12 · Kexin Yi, Jiajun Wu, Chuang Gan, Antonio Torralba, Pushmeet Kohli, Joshua B. Tenenbaum

We marry two powerful ideas: deep representation learning for visual recognition and language understanding, and symbolic program execution for reasoning. Our neural-symbolic visual question answering (NS-VQA) system first recovers a structural scene representation from the image and a program trace from the question. It then executes the program on the scene representation to obtain an answer. Incorporating symbolic structure as prior knowledge offers three unique advantages. First, executing programs on a symbolic space is more robust to long program traces; our model can solve complex reasoning tasks better, achieving an accuracy of 99.8% on the CLEVR dataset. Second, the model is more data- and memory-efficient: it performs well after learning on a small number of training data; it can also encode an image into a compact representation, requiring less storage than existing methods for offline question answering. Third, symbolic program execution offers full transparency to the reasoning process; we are thus able to interpret and diagnose each execution step.

📄 PDF Abstract BibTeX arXiv:1810.02338

Code (2)

kexinyi/ns-vqa pytorch
nerdimite/neuro-symbolic-ai-soc pytorch

Tasks

Question AnsweringRepresentation LearningVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Disentangling Extraction and Reasoning in Multi-hop Spatial Reasoning

2023-10-25 · Roshanak Mirzaee, Parisa Kordjamshidi

Spatial reasoning over text is challenging as the models not only need to extract the direct spatial information from the text but also reason over those and infer implicit spatial relations. Recent studies highlight the…

Spatial Reasoning

UFAL-CUNI at SemEval-2026 Task 11: An Efficient Modular Neuro-symbolic Method for Syllogistic Reasoning

2026-05-06 · Ivan Kartáč, Kristýna Onderková, Jan Bronec, Zdeněk Kasner 외 arxiv

This paper describes our system submitted to SemEval-2026 Task 11: Disentangling Content and Formal Reasoning in Large Language Models. We present an efficient modular neuro-symbolic approach, combining a symbolic prover…

Machine Translation

SEF-CLGC at SemEval-2026 Task 11: Logical Notation Impact on Language Model Performance

2026-06-08 · Hanna Abi Akl, Fabien Gandon, Catherine Faron, Pierre Monnin arxiv

This paper revisits our pipeline called Syllogistic Evaluation Framework-Common Logic Grammar Construction (SEF-CLGC). We combine formal logical notations with Small Language Models (SLMs) to evaluate reasoning performan…

SymbolicThought: Integrating Language Models and Symbolic Reasoning for Consistent and Interpretable Human Relationship Understanding

2025-07-05 · Runcong Zhao, Qinglin Zhu, Hainiu Xu, Bin Liang 외 arxiv

Understanding character relationships is essential for interpreting complex narratives and conducting socially grounded AI research. However, manual annotation is time-consuming and low in coverage, while large language …

Augmented Vision-Language Models: A Systematic Review

2025-07-24 · Anthony C Davis, Burhan Sadiq, Tianmin Shu, Chien-Ming Huang arxiv

Recent advances in visual-language machine learning models have demonstrated exceptional ability to use natural language and understand visual scenes by training on large, unstructured datasets. However, this training pa…

Logical Reasoning