paper-with-me

홈 › Papers

Neural Self Talk: Image Understanding via Continuous Questioning and Answering

2015-12-10 · Yezhou Yang, Yi Li, Cornelia Fermuller, Yiannis Aloimonos

In this paper we consider the problem of continuously discovering image contents by actively asking image based questions and subsequently answering the questions being asked. The key components include a Visual Question Generation (VQG) module and a Visual Question Answering module, in which Recurrent Neural Networks (RNN) and Convolutional Neural Network (CNN) are used. Given a dataset that contains images, questions and their answers, both modules are trained at the same time, with the difference being VQG uses the images as input and the corresponding questions as output, while VQA uses images and questions as input and the corresponding answers as output. We evaluate the self talk process subjectively using Amazon Mechanical Turk, which show effectiveness of the proposed method.

📄 PDF Abstract BibTeX arXiv:1512.03460

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringQuestion GenerationQuestion-GenerationVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

SQ-LLaVA: Self-Questioning for Large Vision-Language Assistant

2024-03-17 · Guohao Sun, Can Qin, Jiamian Wang, Zeyuan Chen 외

Recent advances in vision-language models have shown notable generalization in broad tasks through visual instruction tuning. However, bridging the gap between the pre-trained vision encoder and the large language models…

Language ModellingQuestion AnsweringSelf-Supervised LearningVisual Question Answering

Introspective Growth: Automatically Advancing LLM Expertise in Technology Judgment

2025-05-18 · Siyang Wu, Honglin Bao, Nadav Kunievsky, James A. Evans

Large language models (LLMs) increasingly demonstrate signs of conceptual understanding, yet much of their internal knowledge remains latent, loosely structured, and difficult to access or evaluate. We propose self-quest…

Diagnostic

EmoTalkingGaussian: Continuous Emotion-conditioned Talking Head Synthesis

2025-02-02 · Junuk Cha, Seongro Yoon, Valeriya Strizhkova, Francois Bremond 외

3D Gaussian splatting-based talking head synthesis has recently gained attention for its ability to render high-fidelity images with real-time inference speed. However, since it is typically trained on only a short video…

Self-Supervised LearningSSIMtext-to-speechText to Speech

Socratic Questioning: Learn to Self-guide Multimodal Reasoning in the Wild

2025-01-06 · Wanpeng Hu, Haodi Liu, Lin Chen, Feng Zhou 외

Complex visual reasoning remains a key challenge today. Typically, the challenge is tackled using methodologies such as Chain of Thought (COT) and visual instruction tuning. However, how to organically combine these two …

HallucinationMultimodal ReasoningQuestion AnsweringVisual Reasoning

Talk Structurally, Act Hierarchically: A Collaborative Framework for LLM Multi-Agent Systems

2025-02-16 · Zhao Wang, Sota Moriyama, Wei-Yao Wang, Briti Gangopadhyay 외

Recent advancements in LLM-based multi-agent (LLM-MA) systems have shown promise, yet significant challenges remain in managing communication and refinement when agents collaborate on complex tasks. In this paper, we pro…

Open-Domain Question AnsweringQuestion AnsweringText Generation