paper-with-me

홈 › Papers

Imagery as Inquiry: Exploring A Multimodal Dataset for Conversational Recommendation

2024-05-23 · Se-eun Yoon, Hyunsik Jeon, Julian McAuley

We introduce a multimodal dataset where users express preferences through images. These images encompass a broad spectrum of visual expressions ranging from landscapes to artistic depictions. Users request recommendations for books or music that evoke similar feelings to those captured in the images, and recommendations are endorsed by the community through upvotes. This dataset supports two recommendation tasks: title generation and multiple-choice selection. Our experiments with large foundation models reveal their limitations in these tasks. Particularly, vision-language models show no significant advantage over language-only counterparts that use descriptions, which we hypothesize is due to underutilized visual capabilities. To better harness these abilities, we propose the chain-of-imagery prompting, which results in notable improvements. We release our code and datasets.

📄 PDF Abstract BibTeX arXiv:2405.14142

Code (0)

등록된 구현이 없습니다.

Tasks

Conversational RecommendationMultiple-choice

Similar Papers 제목 키워드 기반

cPAPERS: A Dataset of Situated and Multimodal Interactive Conversations in Scientific Papers

2024-06-12 · Anirudh Sundar, Jin Xu, William Gay, Christopher Richardson 외

An emerging area of research in situated and multimodal interactive conversations (SIMMC) includes interactions in scientific papers. Since scientific papers are primarily composed of text, equations, figures, and tables…

FLAIR-HUB: Large-scale Multimodal Dataset for Land Cover and Crop Mapping

2025-06-08 · Anatol Garioud, Sébastien Giordano, Nicolas David, Nicolas Gonthier

The growing availability of high-quality Earth Observation (EO) data enables accurate global land cover and crop type monitoring. However, the volume and heterogeneity of these datasets pose major processing and annotati…

Earth ObservationMulti-Task Learning

Using Machine Mental Imagery for Representing Common Ground in Situated Dialogue

2026-04-22 · Biswesh Mohapatra, Giovanni Duca, Laurent Romary, Justine Cassell arxiv

Situated dialogue requires speakers to maintain a reliable representation of shared context rather than reasoning only over isolated utterances. Current conversational agents often struggle with this requirement, especia…

Response Generation

Kaleidoscope Gallery: Exploring Ethics and Generative AI Through Art

2025-05-20 · Alayt Issak, Uttkarsh Narayan, Ramya Srinivasan, Erica Kleinman 외

Ethical theories and Generative AI (GenAI) models are dynamic concepts subject to continuous evolution. This paper investigates the visualization of ethics through a subset of GenAI models. We expand on the emerging fiel…

Ethics

Thinking Like a Botanist: Challenging Multimodal Language Models with Intent-Driven Chain-of-Inquiry

2026-04-22 · Syed Nazmus Sakib, Nafiul Haque, Shahrear Bin Amin, Hasan Muhammad Abdullah 외 arxiv

Vision evaluations are typically done through multi-step processes. In most contemporary fields, experts analyze images using structured, evidence-based adaptive questioning. In plant pathology, botanists inspect leaf im…

Question AnsweringVisual GroundingVisual Reasoning