paper-with-me

Papers

An animated picture says at least a thousand words: Selecting Gif-based Replies in Multimodal Dialog

2021-09-24 · Findings (EMNLP) 2021 11 · Xingyao Wang, David Jurgens

Online conversations include more than just text. Increasingly, image-based responses such as memes and animated gifs serve as culturally recognized and often humorous responses in conversation. However, while NLP has broadened to multimodal models, conversational dialog systems have largely focused only on generating text replies. Here, we introduce a new dataset of 1.56M text-gif conversation turns and introduce a new multimodal conversational model Pepe the King Prawn for selecting gif-based replies. We demonstrate that our model produces relevant and high-quality gif responses and, in a large randomized control trial of multiple models replying to real users, we show that our model replies with gifs that are significantly better received by the community.

📄 PDF Abstract BibTeX arXiv:2109.12212

Code (1)

xingyaoww/gif-reply 공식 구현 pytorch

Tasks

Multimodal GIF Dialog

Similar Papers 제목 키워드 기반

Image and Information

2016-02-03 · Frank Nielsen

A well-known old adage says that {\em "A picture is worth a thousand words!"} (attributed to the Chinese philosopher Confucius ca 500 years BC). But more precisely, what do we mean by information in images? And how can i…

MUMU: Bootstrapping Multimodal Image Generation from Text-to-Image Data

2024-06-26 · William Berman, Alexander Peysakhovich

We train a model to generate images from multimodal prompts of interleaved text and images such as "a <picture of a man> man and his <picture of a dog> dog in an <picture of a cartoon> animated style." We bootstrap a mul…

DecoderGPUImage CaptioningImage Generation+3

One Picture is Worth a Thousand Words: A New Wallet Recovery Process

2022-05-05 · Hervé Chabannne, Vincent Despiegel, Linda Guiga

We introduce a new wallet recovery process. Our solution associates 1) visual passwords: a photograph of a secretly picked object (Chabanne et al., 2013) with 2) ImageNet classifiers transforming images into binary vecto…

Retrieval

A Thousand Words Are Worth More Than a Picture: Natural Language-Centric Outside-Knowledge Visual Question Answering

2022-01-14 · Feng Gao, Qing Ping, Govind Thattai, Aishwarya Reganti 외

Outside-knowledge visual question answering (OK-VQA) requires the agent to comprehend the image, make use of relevant knowledge from the entire web, and digest all the information to answer the question. Most previous wo…

Generative Question AnsweringImage to textPassage RetrievalQuestion Answering+3

How Large Language Models Are Changing MOOC Essay Answers: A Comparison of Pre- and Post-LLM Responses

2025-04-17 · Leo Leppänen, Lili Aunimo, Arto Hellas, Jukka K. Nurminen 외

The release of ChatGPT in late 2022 caused a flurry of activity and concern in the academic and educational communities. Some see the tool's ability to generate human-like text that passes at least cursory inspections fo…

Dynamic Topic ModelingEthicsInformation Retrieval