paper-with-me

Papers

Generating by Understanding: Neural Visual Generation with Logical Symbol Groundings

2023-10-26 · Yifei Peng, Zijie Zha, Yu Jin, Zhexu Luo, Wang-Zhou Dai, Zhong Ren, Yao-Xiang Ding, Kun Zhou

Making neural visual generative models controllable by logical reasoning systems is promising for improving faithfulness, transparency, and generalizability. We propose the Abductive visual Generation (AbdGen) approach to build such logic-integrated models. A vector-quantized symbol grounding mechanism and the corresponding disentanglement training method are introduced to enhance the controllability of logical symbols over generation. Furthermore, we propose two logical abduction methods to make our approach require few labeled training data and support the induction of latent logical generative rules from data. We experimentally show that our approach can be utilized to integrate various neural generative models with logical reasoning systems, by both learning from scratch or utilizing pre-trained models directly. The code is released at https://github.com/future-item/AbdGen.

📄 PDF Abstract BibTeX arXiv:2310.17451

Code (2)

candytalking/abdgen 공식 구현 pytorch
future-item/abdgen 공식 구현 pytorch

Tasks

DisentanglementLogical Reasoning

Similar Papers 제목 키워드 기반

V-LoL: A Diagnostic Dataset for Visual Logical Learning

2023-06-13 · Lukas Helff, Wolfgang Stammer, Hikaru Shindo, Devendra Singh Dhami 외

Despite the successes of recent developments in visual AI, different shortcomings still exist; from missing exact logical reasoning, to abstract generalization abilities, to understanding complex and noisy scenes. Unfort…

DiagnosticLogical ReasoningVisual Reasoning

Logically Consistent Language Models via Neuro-Symbolic Integration

2024-09-09 · Diego Calanzone, Stefano Teso, Antonio Vergari

Large language models (LLMs) are a promising venue for natural language understanding and generation. However, current LLMs are far from reliable: they are prone to generating non-factual information and, more crucially,…

Natural Language Understanding

GenEscape: Hierarchical Multi-Agent Generation of Escape Room Puzzles

2025-06-27 · Mengyi Shan, Brian Curless, Ira Kemelmacher-Shlizerman, Steve Seitz

We challenge text-to-image models with generating escape room puzzle images that are visually appealing, logically solid, and intellectually stimulating. While base image models struggle with spatial relationships and af…

Symbol-LLM: Leverage Language Models for Symbolic System in Visual Human Activity Reasoning

2023-11-29 · NeurIPS 2023 11 · Xiaoqian Wu, Yong-Lu Li, Jianhua Sun, Cewu Lu

Human reasoning can be understood as a cooperation between the intuitive, associative "System-1" and the deliberative, logical "System-2". For existing System-1-like methods in visual activity understanding, it is crucia…

Bringing The Consistency Gap: Explicit Structured Memory for Interleaved Image-Text Generation

2025-10-13 · Zeteng Lin, Xingxing Li, Wen You, Xiaoyang Li 외 arxiv

Existing Vision Language Models (VLMs) often struggle to preserve logic, entity identity, and artistic style during extended, interleaved image-text interactions. We identify this limitation as "Multimodal Context Drift"…

multimodal generationText Generation