paper-with-me

홈 › Papers

Sketch2Code: Evaluating Vision-Language Models for Interactive Web Design Prototyping

2024-10-21 · Ryan Li, Yanzhe Zhang, Diyi Yang

Sketches are a natural and accessible medium for UI designers to conceptualize early-stage ideas. However, existing research on UI/UX automation often requires high-fidelity inputs like Figma designs or detailed screenshots, limiting accessibility and impeding efficient design iteration. To bridge this gap, we introduce Sketch2Code, a benchmark that evaluates state-of-the-art Vision Language Models (VLMs) on automating the conversion of rudimentary sketches into webpage prototypes. Beyond end-to-end benchmarking, Sketch2Code supports interactive agent evaluation that mimics real-world design workflows, where a VLM-based agent iteratively refines its generations by communicating with a simulated user, either passively receiving feedback instructions or proactively asking clarification questions. We comprehensively analyze ten commercial and open-source models, showing that Sketch2Code is challenging for existing VLMs; even the most capable models struggle to accurately interpret sketches and formulate effective questions that lead to steady improvement. Nevertheless, a user study with UI/UX experts reveals a significant preference for proactive question-asking over passive feedback reception, highlighting the need to develop more effective paradigms for multi-turn conversational agents.

📄 PDF Abstract BibTeX arXiv:2410.16232

Code (0)

등록된 구현이 없습니다.

Tasks

Benchmarking

Similar Papers 제목 키워드 기반

Interactive Sketchpad: A Multimodal Tutoring System for Collaborative, Visual Problem-Solving

2025-02-12 · Steven-Shine Chen, JiMin Lee, Paul Pu Liang

Humans have long relied on visual aids like sketches and diagrams to support reasoning and problem-solving. Visual tools, like auxiliary lines in geometry or graphs in calculus, are essential for understanding complex id…

Mathmultimodal interaction

Emergent Communication in Interactive Sketch Question Answering

2023-09-21 · NeurIPS 2023 11

Vision-based emergent communication (EC) aims to learn to communicate through sketches and demystify the evolution of human communication. Ironically, previous works neglect multi-round interaction, which is indispensabl…

Sketchforme: Composing Sketched Scenes from Text Descriptions for Interactive Applications

2019-04-08 · Forrest Huang, John F. Canny

Sketching and natural languages are effective communication media for interactive applications. We introduce Sketchforme, the first neural-network-based system that can generate sketches based on text descriptions specif…

SketchVLM: Vision language models can annotate images to explain thoughts and guide users

2026-04-23 · Brandon Collins, Logan Bolton, Hung Huy Nguyen, Mohammad Reza Taesiri 외 arxiv

When answering questions about images, humans naturally point, label, and draw to explain their reasoning. In contrast, modern vision-language models (VLMs) such as Gemini-3-Pro and GPT-5 only respond with text, which ca…

Trajectory PredictionVisual ReasoningObject Counting

Language-based Colorization of Scene Sketches

2019-11-17 · Transactions on Graphics 2019 11 · Changqing Zou, Haoran Mo, Chengying Gao, Ruofei Du 외

Being natural, touchless, and fun-embracing, language-based inputs have been demonstrated effective for various tasks from image generation to literacy education for children. This paper for the first time presents a lan…

ColorizationImage GenerationScene UnderstandingSketch