paper-with-me

홈 › Papers

Visual question answering based evaluation metrics for text-to-image generation

2024-11-15 · Mizuki Miyamoto, Ryugo Morita, Jinjia Zhou

Text-to-image generation and text-guided image manipulation have received considerable attention in the field of image generation tasks. However, the mainstream evaluation methods for these tasks have difficulty in evaluating whether all the information from the input text is accurately reflected in the generated images, and they mainly focus on evaluating the overall alignment between the input text and the generated images. This paper proposes new evaluation metrics that assess the alignment between input text and generated images for every individual object. Firstly, according to the input text, chatGPT is utilized to produce questions for the generated images. After that, we use Visual Question Answering(VQA) to measure the relevance of the generated images to the input text, which allows for a more detailed evaluation of the alignment compared to existing methods. In addition, we use Non-Reference Image Quality Assessment(NR-IQA) to evaluate not only the text-image alignment but also the quality of the generated images. Experimental results show that our proposed evaluation approach is the superior metric that can simultaneously assess finer text-image alignment and image quality while allowing for the adjustment of these ratios.

📄 PDF Abstract BibTeX arXiv:2411.10183

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationImage ManipulationImage Quality AssessmentNR-IQAQuestion AnsweringText to Image GenerationText-to-Image GenerationVisual Question AnsweringVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Focus 설명 없음

Similar Papers 제목 키워드 기반

Fully Authentic Visual Question Answering Dataset from Online Communities

2023-11-27 · Chongyan Chen, Mengchen Liu, Noel Codella, Yunsheng Li 외

Visual Question Answering (VQA) entails answering questions about images. We introduce the first VQA dataset in which all contents originate from an authentic use case. Sourced from online question answering community fo…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

CUS-QA: Local-Knowledge-Oriented Open-Ended Question Answering Dataset

2025-07-30 · Jindřich Libovický, Jindřich Helcl, Andrei Manea, Gianluca Vico arxiv

We introduce CUS-QA, a benchmark for evaluation of open-ended regional question answering that encompasses both textual and visual modalities. We also provide strong baselines using state-of-the-art large language models…

Question Answering

Faithful Multimodal Explanation for Visual Question Answering

2018-09-08 · WS 2019 8 · Jialin Wu, Raymond J. Mooney

AI systems' ability to explain their reasoning is critical to their utility and trustworthiness. Deep neural networks have enabled significant progress on many challenging problems such as visual question answering (VQA)…

Explanatory Visual Question AnsweringQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

AutoViVQA: A Large-Scale Automatically Constructed Dataset for Vietnamese Visual Question Answering

2026-03-10 · Nguyen Anh Tuong, Phan Ba Duc, Nguyen Trung Quoc, Tran Dac Thinh 외 arxiv

Visual Question Answering (VQA) is a fundamental multimodal task that requires models to jointly understand visual and textual information. Early VQA systems relied heavily on language biases, motivating subsequent work …

Visual Question AnsweringRepresentation LearningMachine TranslationImage Captioning

$M^3 QuestionIng$: Multi-modal Multi-span Medical Question Answering

2026-05-19 · Anisha Saha, Vaibhav Rathore, Abhisek Tiwari, Akash Ghosh 외 arxiv

The growing adoption of AI in healthcare, particularly in preventive care, highlights the critical need for accessibility and precision in Medical Question Answering (MedQA). In recent years, significant efforts have bee…

Question Answering