paper-with-me

Papers

Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

2024-03-29 · Weifeng Lin, Xinyu Wei, Ruichuan An, Peng Gao, Bocheng Zou, Yulin Luo, Siyuan Huang, Shanghang Zhang, Hongsheng Li

The interaction between humans and artificial intelligence (AI) is a crucial factor that reflects the effectiveness of multimodal large language models (MLLMs). However, current MLLMs primarily focus on image-level comprehension and limit interaction to textual instructions, thereby constraining their flexibility in usage and depth of response. In this paper, we introduce the Draw-and-Understand project: a new model, a multi-domain dataset, and a challenging benchmark for visual prompting. Specifically, we propose SPHINX-V, a new end-to-end trained Multimodal Large Language Model (MLLM) that connects a vision encoder, a visual prompt encoder and an LLM for various visual prompts (points, bounding boxes, and free-form shape) and language understanding. To advance visual prompting research for MLLMs, we introduce MDVP-Data and MDVP-Bench. MDVP-Data features a multi-domain dataset containing 1.6M unique image-visual prompt-text instruction-following samples, including natural images, document images, OCR images, mobile screenshots, web screenshots, and multi-panel images. Furthermore, we present MDVP-Bench, a comprehensive and challenging benchmark to assess a model's capability in understanding visual prompting instructions. Our experiments demonstrate SPHINX-V's impressive multimodal interaction capabilities through visual prompting, revealing significant improvements in detailed pixel-level description and question-answering abilities.

📄 PDF Abstract BibTeX arXiv:2403.20271

Code (1)

AFeng-x/Draw-and-Understand 공식 구현 pytorch

Tasks

Instruction FollowingLanguage ModellingLarge Language Modelmultimodal interactionMultimodal Large Language ModelOptical Character Recognition (OCR)Question AnsweringVisual Prompting

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Draw an Audio: Leveraging Multi-Instruction for Video-to-Audio Synthesis

2024-09-10 · Qi Yang, Binjie Mao, Zili Wang, Xing Nie 외

Foley is a term commonly used in filmmaking, referring to the addition of daily sound effects to silent films or videos to enhance the auditory experience. Video-to-Audio (V2A), as a particular type of automatic foley ta…

Audio SynthesisAudio-Visual Synchronization

CLIPDrawX: Primitive-based Explanations for Text Guided Sketch Synthesis

2023-12-04 · Nityanand Mathur, Shyam Marjit, Abhra Chaudhuri, Anjan Dutta

With the goal of understanding the visual concepts that CLIP associates with text prompts, we show that the latent space of CLIP can be visualized solely in terms of linear transformations on simple geometric primitives …

MS2Mesh-XR: Multi-modal Sketch-to-Mesh Generation in XR Environments

2024-12-12 · Yuqi Tong, Yue Qiu, Ruiyang Li, Shi Qiu 외

We present MS2Mesh-XR, a novel multi-modal sketch-to-mesh generation pipeline that enables users to create realistic 3D objects in extended reality (XR) environments using hand-drawn sketches assisted by voice inputs. In…

Test-time Correction with Human Feedback: An Online 3D Detection System via Visual Prompting

2024-12-10 · Zetong Yang, Hanxue Zhang, Yanan sun, Li Chen 외

This paper introduces Test-time Correction (TTC) system, a novel online 3D detection system designated for online correction of test-time errors via human feedback, to guarantee the safety of deployed autonomous driving …

Autonomous DrivingVisual Prompting

GenExam: A Multidisciplinary Text-to-Image Exam

2025-09-17 · Zhaokai Wang, Penghao Yin, Xiangyu Zhao, Changyao Tian 외 arxiv

Exams are a fundamental test of expert-level intelligence and require integrated understanding, reasoning, and generation. Existing exam-style benchmarks mainly focus on understanding and reasoning tasks, and current gen…

Image Generation