paper-with-me

홈 › Papers

Several questions of visual generation in 2024

2024-07-11 · Shuyang Gu

This paper does not propose any new algorithms but instead outlines various problems in the field of visual generation based on the author's personal understanding. The core of these problems lies in how to decompose visual signals, with all other issues being closely related to this central problem and stemming from unsuitable approaches to signal decomposition. This paper aims to draw researchers' attention to the significance of Visual Signal Decomposition.

📄 PDF Abstract BibTeX arXiv:2407.18290

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

Unified Questioner Transformer for Descriptive Question Generation in Goal-Oriented Visual Dialogue

2021-06-29 · ICCV 2021 10 · Shoya Matsumori, Kosuke Shingyouchi, Yuki Abe, Yosuke Fukuchi 외

Building an interactive artificial intelligence that can ask questions about the real world is one of the biggest challenges for vision and language problems. In particular, goal-oriented visual dialogue, where the aim o…

DescriptiveQuestion GenerationQuestion-Generation

Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries

2025-02-23 · Yin Wu, Quanyu Long, Jing Li, Jianfei Yu 외

Retrieval-Augmented Generation (RAG) is a popular approach for enhancing Large Language Models (LLMs) by addressing their limitations in verifying facts and answering knowledge-intensive questions. As the research in LLM…

BenchmarkingImage RetrievalQuestion AnsweringRAG+2

What You See is What You Ask: Evaluating Audio Descriptions

2025-10-01 · Divy Kala, Eshika Khandelwal, Makarand Tapaswi arxiv

Audio descriptions (ADs) narrate important visual details in movies, enabling Blind and Low Vision (BLV) users to understand narratives and appreciate visual details. Existing works in automatic AD generation mostly focu…

Visual Curiosity: Learning to Ask Questions to Learn Visual Recognition

2018-10-01 · Jianwei Yang, Jiasen Lu, Stefan Lee, Dhruv Batra 외

In an open-world setting, it is inevitable that an intelligent agent (e.g., a robot) will encounter visual objects, attributes or relationships it does not recognize. In this work, we develop an agent empowered with visu…

Question GenerationQuestion-GenerationReinforcement Learning

Automatic Generation of Grounded Visual Questions

2016-12-20 · Shijie Zhang, Lizhen Qu, ShaoDi You, Zhenglu Yang 외

In this paper, we propose the first model to be able to generate visually grounded questions with diverse types for a single image. Visual question generation is an emerging topic which aims to ask questions in natural l…

DiversityQuestion GenerationQuestion-Generation