Underspecification in Scene Description-to-Depiction Tasks
Questions regarding implicitness, ambiguity and underspecification are crucial for understanding the task validity and ethical concerns of multimodal image+text systems, yet have received little attention to date. This position paper maps out a conceptual framework to address this gap, focusing on systems which generate images depicting scenes from scene descriptions. In doing so, we account for how texts and images convey meaning differently. We outline a set of core challenges concerning textual and visual ambiguity, as well as risks that may be amplified by ambiguous and underspecified elements. We propose and discuss strategies for addressing these challenges, including generating visually ambiguous images, and generating a set of diverse images.
Code (0)
등록된 구현이 없습니다.
Tasks
PositionSimilar Papers 제목 키워드 기반
Visual Conceptual Blending with Large-scale Language and Vision Models
We ask the question: to what extent can recent large-scale language and image generation models blend visual concepts? Given an arbitrary object, we identify a relevant object and generate a single-sentence description o…
Image GenerationLanguage ModelingLanguage ModellingObject+1Novel Artistic Scene-Centric Datasets for Effective Transfer Learning in Fragrant Spaces
Olfaction, often overlooked in cultural heritage studies, holds profound significance in shaping human experiences and identities. Examining historical depictions of olfactory scenes can offer valuable insights into the …
Scene ClassificationTransfer LearningFairness and underspecification in acoustic scene classification: The case for disaggregated evaluations
Underspecification and fairness in machine learning (ML) applications have recently become two prominent issues in the ML community. Acoustic scene classification (ASC) applications have so far remained unaffected by thi…
Acoustic Scene ClassificationFairnessScene ClassificationUnderspecification in Language Modeling Tasks: A Causality-Informed Study of Gendered Pronoun Resolution
Modern language modeling tasks are often underspecified: for a given token prediction, many words may satisfy the user's intent of producing natural language at inference time, however only one word will minimize the tas…
Language ModelingLanguage ModellingSelection biasCLIPascene: Scene Sketching with Different Types and Levels of Abstraction
In this paper, we present a method for converting a given scene image into a sketch using different types and multiple levels of abstraction. We distinguish between two types of abstraction. The first considers the fidel…
Disentanglement