paper-with-me

홈 › Papers

Hallucination as an Upper Bound: A New Perspective on Text-to-Image Evaluation

2025-09-25 · Seyed Amir Kasaei, Mohammad Hossein Rohban arxiv

In language and vision-language models, hallucination is broadly understood as content generated from a model's prior knowledge or biases rather than from the given input. While this phenomenon has been studied in those domains, it has not been clearly framed for text-to-image (T2I) generative models. Existing evaluations mainly focus on alignment, checking whether prompt-specified elements appear, but overlook what the model generates beyond the prompt. We argue for defining hallucination in T2I as bias-driven deviations and propose a taxonomy with three categories: attribute, relation, and object hallucinations. This framing introduces an upper bound for evaluation and surfaces hidden biases, providing a foundation for richer assessment of T2I models.

📄 PDF Abstract BibTeX arXiv:2509.21257

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Hallucination as Exploit: Evidence-Carrying Multimodal Agents

2026-05-18 · Guijia Zhang, Hao Zheng, Harry Yang arxiv

Multimodal agents increasingly choose tool calls from screenshots, documents, and webpages, where a false perceptual claim can turn hallucination from an answer-quality error into an authorization failure. We formalize t…

Revisiting Prioritized Experience Replay: A Value Perspective

2021-02-05 · Ang A. Li, Zongqing Lu, Chenglin Miao

Experience replay enables off-policy reinforcement learning (RL) agents to utilize past experiences to maximize the cumulative reward. Prioritized experience replay that weighs experiences by the magnitude of their tempo…

Atari GamesQ-LearningReinforcement Learning (RL)

Towards Lossless Implicit Neural Representation via Bit Plane Decomposition

2025-02-28 · CVPR 2025 1 · Woo Kyoung Han, Byeonghun Lee, Hyunmin Cho, Sunghoon Im 외

We quantify the upper bound on the size of the implicit neural representation (INR) model from a digital perspective. The upper bound of the model size increases exponentially as the required bit-precision increases. To …

Image CompressionQuantization

ChainMPQ: Interleaved Text-Image Reasoning Chains for Mitigating Relation Hallucinations

2025-10-07 · Yike Wu, Yiwei Wang, Yujun Cai arxiv

While Large Vision-Language Models (LVLMs) achieve strong performance in multimodal tasks, hallucinations continue to hinder their reliability. Among the three categories of hallucinations, which include object, attribut…

Relational Reasoning

Sharp Bounds for Federated Averaging (Local SGD) and Continuous Perspective

2021-11-05 · Margalit Glasgow, Honglin Yuan, Tengyu Ma

Federated Averaging (FedAvg), also known as Local SGD, is one of the most popular algorithms in Federated Learning (FL). Despite its simplicity and popularity, the convergence rate of FedAvg has thus far been undetermine…

Federated Learning