paper-with-me

홈 › Papers

Multi-modal dialog for browsing large visual catalogs using exploration-exploitation paradigm in a joint embedding space

2019-01-28 · Indrani Bhattacharya, Arkabandhu Chowdhury, Vikas Raykar

We present a multi-modal dialog system to assist online shoppers in visually browsing through large catalogs. Visual browsing is different from visual search in that it allows the user to explore the wide range of products in a catalog, beyond the exact search matches. We focus on a slightly asymmetric version of the complete multi-modal dialog where the system can understand both text and image queries but responds only in images. We formulate our problem of "showing $k$ best images to a user" based on the dialog context so far, as sampling from a Gaussian Mixture Model in a high dimensional joint multi-modal embedding space, that embed both the text and the image queries. Our system remembers the context of the dialog and uses an exploration-exploitation paradigm to assist in visual browsing. We train and evaluate the system on a multi-modal dialog dataset that we generate from large catalog data. Our experiments are promising and show that the agent is capable of learning and can display relevant results with an average cosine similarity of 0.85 to the ground truth. Our preliminary human evaluation also corroborates the fact that such a multi-modal dialog system for visual browsing is well-received and is capable of engaging human users.

📄 PDF Abstract BibTeX arXiv:1901.09854

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

BrowseComp-$V^3$: A Visual, Vertical, and Verifiable Benchmark for Multimodal Browsing Agents

2026-02-13 · Huanyao Zhang, Jiepeng Zhou, Bo Li, Bowen Zhou 외 arxiv

Multimodal large language models (MLLMs), equipped with increasingly advanced planning and tool-use capabilities, are evolving into autonomous agents capable of performing multimodal web browsing and deep search in open-…

VisBrowse-Bench: Benchmarking Visual-Native Search for Multimodal Browsing Agents

2026-03-17 · Zhengbo Zhang, Jinbo Su, Zhaowen Zhou, Changtao Miao 외 arxiv

The rapid advancement of Multimodal Large Language Models (MLLMs) has enabled browsing agents to acquire and reason over multimodal information in the real world. But existing benchmarks suffer from two limitations: insu…

Visual ReasoningImage Retrieval

MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents

2025-08-14 · Shilong Li, Xingyuan Bu, Wenjie Wang, Jiaheng Liu 외 arxiv

AI agents with advanced reasoning and tool-use capabilities have demonstrated impressive performance in web browsing for deep search. However, existing benchmarks such as BrowseComp primarily focus on textual content, ov…

Multimodal Reasoning

OpenViDial 2.0: A Larger-Scale, Open-Domain Dialogue Generation Dataset with Visual Contexts

2021-09-27 · Shuhe Wang, Yuxian Meng, Xiaoya Li, Xiaofei Sun 외

In order to better simulate the real human conversation process, models need to generate dialogue utterances based on not only preceding textual contexts but also visual contexts. However, with the development of multi-m…

Dialogue GenerationMulti-modal Dialogue Generation

MMSearch-Plus: Benchmarking Provenance-Aware Search for Multimodal Browsing Agents

2025-08-29 · Xijia Tao, Yihua Teng, Xinxing Su, Xinyu Fu 외 arxiv

Existing multimodal browsing benchmarks often fail to require genuine multimodal reasoning, as many tasks can be solved with text-only heuristics without vision-in-the-loop verification. We introduce MMSearch-Plus, a 311…

Multimodal ReasoningText Retrieval