paper-with-me

Papers

What Do AI-Generated Images Want?

2025-10-23 · Amanda Wasielewski arxiv

W.J.T. Mitchell's influential essay 'What do pictures want?' shifts the theoretical focus away from the interpretative act of understanding pictures and from the motivations of the humans who create them to the possibility that the picture itself is an entity with agency and wants. In this article, I reframe Mitchell's question in light of contemporary AI image generation tools to ask: what do AI-generated images want? Drawing from art historical discourse on the nature of abstraction, I argue that AI-generated images want specificity and concreteness because they are fundamentally abstract. Multimodal text-to-image models, which are the primary subject of this article, are based on the premise that text and image are interchangeable or exchangeable tokens and that there is a commensurability between them, at least as represented mathematically in data. The user pipeline that sees textual input become visual output, however, obscures this representational regress and makes it seem like one form transforms into the other -- as if by magic.

📄 PDF Abstract BibTeX arXiv:2510.20350

Code (0)

등록된 구현이 없습니다.

Tasks

Image Generation

Similar Papers 제목 키워드 기반

Text to artistic image generation

2022-05-05 · Qinghe Tian, Jean-Claude Franchitti

Painting is one of the ways for people to express their ideas, but what if people with disabilities in hands want to paint? To tackle this challenge, we create an end-to-end solution that can generate artistic images fro…

Generative Adversarial NetworkImage Generation

Customizing Language Model Responses with Contrastive In-Context Learning

2024-01-30 · Xiang Gao, Kamalika Das

Large language models (LLMs) are becoming increasingly important for machine learning applications. However, it can be challenging to align LLMs with our intent, particularly when we want to generate content that is pref…

In-Context LearningLanguage ModelingLanguage Modelling

Get What You Want, Not What You Don't: Image Content Suppression for Text-to-Image Diffusion Models

2024-02-08 · Senmao Li, Joost Van de Weijer, Taihang Hu, Fahad Shahbaz Khan 외

The success of recent text-to-image diffusion models is largely due to their capacity to be guided by a complex text prompt, which enables users to precisely describe the desired content. However, these models struggle t…

Who’s Doing What: Joint Modeling of Names and Verbs for Simultaneous Face and Pose Annotation

2009-12-01 · NeurIPS 2009 12 · Jie Luo, Barbara Caputo, Vittorio Ferrari

Given a corpus of news items consisting of images accompanied by text captions, we want to find out ``whos doing what, i.e. associate names and action verbs in the captions to the face and body pose of the persons in the…

Do Automatic Factuality Metrics Measure Factuality? A Critical Evaluation

2024-11-25 · Sanjana Ramprasad, Byron C. Wallace

Modern LLMs can now produce highly readable abstractive summaries, to the point where traditional automated metrics for evaluating summary quality, such as ROUGE, have become saturated. However, LLMs still sometimes intr…