paper-with-me

홈 › Papers

Learning Words by Drawing Images

2019-06-01 · CVPR 2019 6 · Didac Suris, Adria Recasens, David Bau, David Harwath, James Glass, Antonio Torralba

We propose a framework for learning through drawing. Our goal is to learn the correspondence between spoken words and abstract visual attributes, from a dataset of spoken descriptions of images. Building upon recent findings that GAN representations can be manipulated to edit semantic concepts in the generated output, we propose a new method to use such GAN-generated images to train a model using a triplet loss. To apply the method, we develop Audio CLEVRGAN, a new dataset of audio descriptions of GAN-generated CLEVR images, and we describe a training procedure that creates a curriculum of GAN-generated images that focuses training on image pairs that differ in a specific, informative way. Training is done without additional supervision beyond the spoken captions and the GAN. We find that training that takes advantage of GAN-generated edited examples results in improvements in the model's ability to learn attributes compared to previous results. Our proposed learning framework also results in models that can associate spoken words with some abstract visual concepts such as color and size.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Triplet

Methods 이 논문이 사용한 방법론

Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Dogecoin Customer Service Number +1-833-534-1729 설명 없음

Similar Papers 제목 키워드 기반

Emotion-Aware Image Generation from Korean Diary Text via LLM-based Prompt Translation and LoRA Fine-Tuning

2026-06-04 · Jihun Cho, Soo-Yeon Jeong, Sun-Young Ihm arxiv

T2I models cannot effectively capture sentiment from various types of text, including diaries, as they primarily focus on visual object-related patterns rather than contextual emotional understanding. This paper proposes…

Image Generation

Towards An Angry-Birds-like Game System for Promoting Mental Well-being of Players Using Art-Therapy-embedded PCG

2019-11-07 · Zhou Fang, Pujana Paliyawan, Ruck Thawonmas, Tomohiro Harada

This paper presents an integration of a game system and the art therapy concept for promoting the mental well-being of video game players. In the proposed game system, the player plays an Angry-Birds-like game in which l…

Iconary: A Pictionary-Based Game for Testing Multimodal Communication with Drawings and Text

2021-12-01 · EMNLP 2021 11 · Christopher Clark, Jordi Salvador, Dustin Schwenk, Derrick Bonafilia 외

Communicating with humans is challenging for AIs because it requires a shared understanding of the world, complex semantics (e.g., metaphors or analogies), and at times multi-modal gestures (e.g., pointing with a finger,…

World Knowledge

IDraw: Artist Verification from Digital Drawing Images

2026-08-03 · Nayoung Kim, Nan Jiang, Bangjie Sun, Jaewon Shin 외 arxiv

As digital drawings are increasingly shared online, reliable authorship verification has become important for protecting artists and resolving disputes. Yet when authorship is questioned, verification may have to rely on…

Directed Diffusion: Direct Control of Object Placement through Attention Guidance

2023-02-25 · Wan-Duo Kurt Ma, J. P. Lewis, Avisek Lahiri, Thomas Leung 외

Text-guided diffusion models such as DALLE-2, Imagen, eDiff-I, and Stable Diffusion are able to generate an effectively endless variety of images given only a short text prompt describing the desired image content. In ma…