paper-with-me

홈 › Papers

The Factuality Tax of Diversity-Intervened Text-to-Image Generation: Benchmark and Fact-Augmented Intervention

2024-06-29 · Yixin Wan, Di wu, Haoran Wang, Kai-Wei Chang

Prompt-based "diversity interventions" are commonly adopted to improve the diversity of Text-to-Image (T2I) models depicting individuals with various racial or gender traits. However, will this strategy result in nonfactual demographic distribution, especially when generating real historical figures. In this work, we propose DemOgraphic FActualIty Representation (DoFaiR), a benchmark to systematically quantify the trade-off between using diversity interventions and preserving demographic factuality in T2I models. DoFaiR consists of 756 meticulously fact-checked test instances to reveal the factuality tax of various diversity prompts through an automated evidence-supported evaluation pipeline. Experiments on DoFaiR unveil that diversity-oriented instructions increase the number of different gender and racial groups in DALLE-3's generations at the cost of historically inaccurate demographic distributions. To resolve this issue, we propose Fact-Augmented Intervention (FAI), which instructs a Large Language Model (LLM) to reflect on verbalized or retrieved factual information about gender and racial compositions of generation subjects in history, and incorporate it into the generation context of T2I models. By orienting model generations using the reflected historical truths, FAI significantly improves the demographic factuality under diversity interventions while preserving diversity.

📄 PDF Abstract BibTeX arXiv:2407.00377

Code (1)

uclanlp/diverse-factual 공식 구현

Tasks

DiversityImage GenerationLanguage ModelingLanguage ModellingLarge Language ModelText to Image GenerationText-to-Image Generation

Similar Papers 제목 키워드 기반

A Factuality and Diversity Reconciled Decoding Method for Knowledge-Grounded Dialogue Generation

2024-07-08 · Chenxu Yang, Zheng Lin, Chong Tian, Liang Pang 외

Grounding external knowledge can enhance the factuality of responses in dialogue generation. However, excessive emphasis on it might result in the lack of engaging and diverse expressions. Through the introduction of ran…

Dialogue GenerationDiversity

Multi-FAct: Assessing Factuality of Multilingual LLMs using FActScore

2024-02-28 · Sheikh Shafayat, Eunsu Kim, Juhyun Oh, Alice Oh

Evaluating the factuality of long-form large language model (LLM)-generated text is an important challenge. Recently there has been a surge of interest in factuality evaluation for English, but little is known about the …

DiversityFormHallucinationLanguage Modeling+3

REAL Sampling: Boosting Factuality and Diversity of Open-Ended Generation via Asymptotic Entropy

2024-06-11 · Haw-Shiuan Chang, Nanyun Peng, Mohit Bansal, Anil Ramakrishna 외

Decoding methods for large language models (LLMs) usually struggle with the tradeoff between ensuring factuality and maintaining diversity. For example, a higher p threshold in the nucleus (top-p) sampling increases the …

DiversityHallucination

T2I-FactualBench: Benchmarking the Factuality of Text-to-Image Models with Knowledge-Intensive Concepts

2024-12-05 · Ziwei Huang, Wanggui He, Quanyu Long, Yandi Wang 외

Evaluating the quality of synthesized images remains a significant challenge in the development of text-to-image (T2I) generation. Most existing studies in this area primarily focus on evaluating text-image alignment, im…

BenchmarkingImage GenerationMemorizationQuestion Answering+4

Fact-or-Fair: A Checklist for Behavioral Testing of AI Models on Fairness-Related Queries

2025-02-09 · Jen-tse Huang, Yuhang Yan, Linqi Liu, Yixin Wan 외

The generation of incorrect images, such as depictions of people of color in Nazi-era uniforms by Gemini, frustrated users and harmed Google's reputation, motivating us to investigate the relationship between accurately …

DiversityFairnessWorld Knowledge