paper-with-me

홈 › Papers

MMCode: Benchmarking Multimodal Large Language Models for Code Generation with Visually Rich Programming Problems

2024-04-15 · Kaixin Li, Yuchen Tian, Qisheng Hu, Ziyang Luo, Zhiyong Huang, Jing Ma

Programming often involves converting detailed and complex specifications into code, a process during which developers typically utilize visual aids to more effectively convey concepts. While recent developments in Large Multimodal Models have demonstrated remarkable abilities in visual reasoning and mathematical tasks, there is little work on investigating whether these models can effectively interpret visual elements for code generation. To this end, we present MMCode, the first multi-modal coding dataset for evaluating algorithmic problem-solving skills in visually rich contexts. MMCode contains 3,548 questions and 6,620 images collected from real-world programming challenges harvested from 10 code competition websites, presenting significant challenges due to the extreme demand for reasoning abilities. Our experiment results show that current state-of-the-art models struggle to solve these problems. The results highlight the lack of powerful vision-code models, and we hope MMCode can serve as an inspiration for future works in this domain. The data and code are publicly available at https://github.com/likaixin2000/MMCode.

📄 PDF Abstract BibTeX arXiv:2404.09486

Code (3)

happylkx/mmcode 공식 구현
likaixin2000/mmcode 공식 구현
yuchen814/codehalu pytorch

Tasks

BenchmarkingCode GenerationVisual Reasoning

Similar Papers 제목 키워드 기반

Code-Vision: Evaluating Multimodal LLMs Logic Understanding and Code Generation Capabilities

2025-02-17 · Hanbin Wang, Xiaoxuan Zhou, Zhipeng Xu, Keyuan Cheng 외

This paper introduces Code-Vision, a benchmark designed to evaluate the logical understanding and code generation capabilities of Multimodal Large Language Models (MLLMs). It challenges MLLMs to generate a correct progra…

Code GenerationHumanEvalMathMathematical Problem-Solving+1

SVRepair: Structured Visual Reasoning for Automated Program Repair

2026-02-05 · Xiaoxuan Tang, Jincheng Wang, Liwei Luo, Jingxuan Xu 외 arxiv

Large language models (LLMs) have recently shown strong potential for Automated Program Repair (APR), yet most existing approaches remain unimodal and fail to leverage the rich diagnostic signals contained in visual arti…

Visual ReasoningProgram Repair

WildScore: Benchmarking MLLMs in-the-Wild Symbolic Music Reasoning

2025-09-05 · Gagan Mundada, Yash Vishe, Amit Namburi, Xin Xu 외 arxiv

Recent advances in Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities across various vision-language tasks. However, their reasoning abilities in the multimodal symbolic music domain remai…

Question Answering

Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering

2024-03-05 · Chenglei Si, Yanzhe Zhang, Ryan Li, Zhengyuan Yang 외

Generative AI has made rapid advancements in recent years, achieving unprecedented capabilities in multimodal understanding and code generation. This can enable a new paradigm of front-end development in which multimodal…

BenchmarkingCode Generation

Dissecting Dissonance: Benchmarking Large Multimodal Models Against Self-Contradictory Instructions

2024-08-02 · Jin Gao, Lei Gan, Yuankai Li, Yixin Ye 외

Large multimodal models (LMMs) excel in adhering to human instructions. However, self-contradictory instructions may arise due to the increasing trend of multimodal interaction and context length, which is challenging fo…

Benchmarkingmultimodal interaction