paper-with-me

홈 › Papers

TurtleBench: A Visual Programming Benchmark in Turtle Geometry

2024-10-31 · Sina Rismanchian, Yasaman Razeghi, Sameer Singh, Shayan Doroudi

Humans have the ability to reason about geometric patterns in images and scenes from a young age. However, developing large multimodal models (LMMs) capable of similar reasoning remains a challenge, highlighting the need for robust evaluation methods to assess these capabilities. We introduce TurtleBench, a benchmark designed to evaluate LMMs' capacity to interpret geometric patterns -- given visual examples, textual instructions, or both -- and generate precise code outputs. Inspired by turtle geometry, a notion used to teach children foundational coding and geometric concepts, TurtleBench features tasks with patterned shapes that have underlying algorithmic logic. Our evaluation reveals that leading LMMs struggle significantly with these tasks, with GPT-4o achieving only 19\% accuracy on the simplest tasks and few-shot prompting only marginally improves their performance ($<2\%$). TurtleBench highlights the gap between human and AI performance in intuitive and visual geometrical understanding, setting the stage for future research in this area. TurtleBench stands as one of the few benchmarks to evaluate the integration of visual understanding and code generation capabilities in LMMs, setting the stage for future research. Code and Dataset for this paper is provided here: https://github.com/sinaris76/TurtleBench

📄 PDF Abstract BibTeX arXiv:2411.00264

Code (1)

sinaris76/turtlebench 공식 구현

Tasks

Code Generation

Similar Papers 제목 키워드 기반

TurtleBench: Evaluating Top Language Models via Real-World Yes/No Puzzles

2024-10-07 · Qingchen Yu, Shichao Song, Ke Fang, Yunfeng Shi 외

As the application of Large Language Models (LLMs) expands, the demand for reliable evaluations increases. Existing LLM evaluation benchmarks primarily rely on static datasets, making it challenging to assess model perfo…

Logical Reasoning

TurtleAI: Benchmarking Multimodal Models for Visual Programming in Turtle Graphics

2026-06-02 · Chao Wen, Jacqueline Staub, Adish Singla arxiv

Vision-language models (VLMs) have been explored for visual programming, where they generate code to solve visual tasks. However, most prior work focuses on visual programming for productivity; it remains unclear how wel…

Spatial ReasoningVisual Reasoning

Let Go of Your Labels with Unsupervised Transfer

2024-06-11 · International Conference on Machine Learning 2024 6 · Artyom Gadetsky, Yulun Jiang, Maria Brbic

Foundation vision-language models have enabled remarkable zero-shot transferability of the pre-trained representations to a wide range of downstream tasks. However, to solve a new task, zero-shot transfer still necessita…

Image ClusteringUnsupervised Image Classification

SeaTurtleID2022: A long-span dataset for reliable sea turtle re-identification

2023-11-09 · Lukáš Adam, Vojtěch Čermák, Kostas Papafitsoros, Lukáš Picek

This paper introduces the first public large-scale, long-span dataset with sea turtle photographs captured in the wild -- SeaTurtleID2022 (https://www.kaggle.com/datasets/wildlifedatasets/seaturtleid2022). The dataset co…

BenchmarkingInstance SegmentationSegmentationSemantic Segmentation

SeaTurtleID2022: A long-span dataset for reliable sea turtle re-identification

2022-11-18 · Lukáš Adam, Vojtěch Čermák, Kostas Papafitsoros, Lukáš Picek

This paper introduces the first public large-scale, long-span dataset with sea turtle photographs captured in the wild -- \href{https://www.kaggle.com/datasets/wildlifedatasets/seaturtleid2022}{SeaTurtleID2022}. The data…

BenchmarkingInstance SegmentationSegmentationSemantic Segmentation