paper-with-me

홈 › Papers

WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation

2025-08-22 · Rabiul Awal, Mahsa Massoud, Aarash Feizi, Zichao Li, Suyuchen Wang, Christopher Pal, Aishwarya Agrawal, David Vazquez, Siva Reddy, Juan A. Rodriguez, Perouz Taslakian, Spandana Gella, Sai Rajeswar arxiv

We present WebMMU, a multilingual benchmark that evaluates three core web tasks: (1) website visual question answering, (2) code editing involving HTML/CSS/JavaScript, and (3) mockup-to-code generation. Unlike prior benchmarks that treat these tasks separately, WebMMU unifies them using expert-annotated, real-world web data to assess models' abilities in complex multi-step reasoning, precise element grounding, and functional UI comprehension and coding. Our evaluation shows that while multimodal large language models (MLLMs) perform well on basic information extraction, they struggle with reasoning and grounding, editing code to preserve functionality, and generating design-to-code that maintains hierarchy and supports multilingual content. These findings reveal key limitations in current MLLMs and underscore the need for improved multimodal and cross-lingual reasoning to build future web agents capable of automating diverse web development tasks.

📄 PDF Abstract BibTeX arXiv:2508.16763

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question AnsweringInformation ExtractionCode Generation

Similar Papers 제목 키워드 기반

MultiNews: A Web collection of an Aligned Multimodal and Multilingual Corpus

2017-11-01 · WS 2017 11 · Haithem Afli, Pintu Lohar, Andy Way

Integrating Natural Language Processing (NLP) and computer vision is a promising effort. However, the applicability of these methods directly depends on the availability of a specific multimodal data that includes images…

ArticlesContent-Based Image RetrievalImage RetrievalMachine Translation+1

EXAM$^2$: $\underline{Ex}tending$ $\underline{A}udio$ $Understanding$ $in$ $\underline{M}ultilingual$ $and$ $\underline{M}ultimodal$ $Analysis$

2026-08-24 · Jiawen Wang, Xiaoxue Gao, Zi Haur Pang, Nancy F. Chen arxiv

Recent large audio language models (LALMs) have achieved impressive progress in audio understanding. However, existing evaluations remain largely constrained to English and narrow audio domains. Prior benchmarks typicall…

XFUND: A Benchmark Dataset for Multilingual Visually Rich Form Understanding

2022-05-01 · Findings (ACL) 2022 5 · Yiheng Xu, Tengchao Lv, Lei Cui, Guoxin Wang 외

Multimodal pre-training with text, layout, and image has achieved SOTA performance for visually rich document understanding tasks recently, which demonstrates the great potential for joint learning across different modal…

document understandingForm

M4U: Evaluating Multilingual Understanding and Reasoning for Large Multimodal Models

2024-05-24 · Hongyu Wang, Jiayu Xu, Senwei Xie, Ruiping Wang 외

Multilingual multimodal reasoning is a core component in achieving human-level intelligence. However, most existing benchmarks for multilingual multimodal reasoning struggle to differentiate between models of varying per…

Multimodal Reasoning

M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models

2023-06-08 · NeurIPS 2023 11 · Wenxuan Zhang, Sharifah Mahani Aljunied, Chang Gao, Yew Ken Chia 외

Despite the existence of various benchmarks for evaluating natural language processing models, we argue that human exams are a more suitable means of evaluating general intelligence for large language models (LLMs), as t…