paper-with-me

홈 › Papers

Scaffolding Coordinates to Promote Vision-Language Coordination in Large Multi-Modal Models

2024-02-19 · Xuanyu Lei, Zonghan Yang, Xinrui Chen, Peng Li, Yang Liu

State-of-the-art Large Multi-Modal Models (LMMs) have demonstrated exceptional capabilities in vision-language tasks. Despite their advanced functionalities, the performances of LMMs are still limited in challenging scenarios that require complex reasoning with multiple levels of visual information. Existing prompting techniques for LMMs focus on either improving textual reasoning or leveraging tools for image preprocessing, lacking a simple and general visual prompting scheme to promote vision-language coordination in LMMs. In this work, we propose Scaffold prompting that scaffolds coordinates to promote vision-language coordination. Specifically, Scaffold overlays a dot matrix within the image as visual information anchors and leverages multi-dimensional coordinates as textual positional references. Extensive experiments on a wide range of challenging vision-language tasks demonstrate the superiority of Scaffold over GPT-4V with the textual CoT prompting. Our code is released in https://github.com/leixy20/Scaffold.

📄 PDF Abstract BibTeX arXiv:2402.12058

Code (1)

leixy20/scaffold 공식 구현

Tasks

Visual Prompting

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Searching for Synergy in Shared Workspace Human-AI Collaboration

2026-06-16 · Nachiket Kotalwar, Rohini Das, Carolyn Rose arxiv

Automated AI agents are increasingly capable, yet many scientific and professional tasks require human judgment and contextual expertise. We use simulated shared-workspace human-AI teams as a controlled testbed for study…

Learning Dynamic Structural Specialization for Underwater Salient Object Detection

2026-05-15 · Lin Hong, Chenhui Wang, Linan Deng, Yuning Cui 외 arxiv

Underwater salient object detection (USOD) has attracted increasing attention for underwater visual scene understanding and vision-guided robotic applications. However, existing USOD methods still struggle with underwate…

Salient Object DetectionObject LocalizationScene Understanding

Fuzzy, Symbolic, and Contextual: Enhancing LLM Instruction via Cognitive Scaffolding

2025-08-28 · Vanessa Figueiredo arxiv

We study how prompt-level inductive biases influence the cognitive behavior of large language models (LLMs) in instructional dialogue. We introduce a symbolic scaffolding method paired with a short-term memory schema des…

MotifBench: A standardized protein design benchmark for motif-scaffolding problems

2025-02-18 · Zhuoqi Zheng, Bo Zhang, Kieran Didi, Kevin K. Yang 외

The motif-scaffolding problem is a central task in computational protein design: Given the coordinates of atoms in a geometry chosen to confer a desired biochemical function (a motif), the task is to identify diverse pro…

Protein DesignProtein Structure Prediction

Inner Speech as Behavior Guides: Steerable Imitation of Diverse Behaviors for Human-AI coordination

2026-02-24 · Rakshit Trivedi, Kartik Sharma, David C Parkes arxiv

Effective human-AI coordination requires artificial agents capable of exhibiting and responding to human-like behaviors while adapting to changing contexts. Imitation learning has emerged as one of the prominent approach…