paper-with-me

홈 › Papers

PCEval: A Benchmark for Evaluating Physical Computing Capabilities of Large Language Models

2025-12-31 · Inpyo Song, Eunji Jeon, Jangwon Lee arxiv

Large Language Models (LLMs) have demonstrated remarkable capabilities across various domains, including software development, education, and technical assistance. Among these, software development is one of the key areas where LLMs are increasingly adopted. However, when hardware constraints are considered-for instance, in physical computing, where software must interact with and control physical hardware -their effectiveness has not been fully explored. To address this gap, we introduce \textsc{PCEval} (Physical Computing Evaluation), the first benchmark in physical computing that enables a fully automatic evaluation of the capabilities of LLM in both the logical and physical aspects of the projects, without requiring human assessment. Our evaluation framework assesses LLMs in generating circuits and producing compatible code across varying levels of project complexity. Through comprehensive testing of 13 leading models, \textsc{PCEval} provides the first reproducible and automatically validated empirical assessment of LLMs' ability to reason about fundamental hardware implementation constraints within a simulation environment. Our findings reveal that while LLMs perform well in code generation and logical circuit design, they struggle significantly with physical breadboard layout creation, particularly in managing proper pin connections and avoiding circuit errors. \textsc{PCEval} advances our understanding of AI assistance in hardware-dependent computing environments and establishes a foundation for developing more effective tools to support physical computing education.

📄 PDF Abstract BibTeX arXiv:2601.02404

Code (0)

등록된 구현이 없습니다.

Tasks

Code Generation

Similar Papers 제목 키워드 기반

MPCEval: A Benchmark for Multi-Party Conversation Generation

2026-03-05 · Minxing Zhang, Yi Yang, Zhuofan Jia, Xuan Yang 외 arxiv

Multi-party conversation generation, such as smart reply and collaborative assistants, is an increasingly important capability of generative AI, yet its evaluation remains a critical bottleneck. Compared to two-party dia…

Reservoir Computing Benchmarks: a tutorial review and critique

2024-05-10 · Chester Wringe, Martin Trefzer, Susan Stepney

Reservoir Computing is an Unconventional Computation model to perform computation on various different substrates, such as recurrent neural networks or physical materials. The method takes a 'black-box' approach, trainin…

OpenPRC: A Unified Open-Source Framework for Physics-to-Task Evaluation in Physical Reservoir Computing

2026-04-08 · Yogesh Phalak, Wen Sin Lor, Apoorva Khairnar, Benjamin Jantzen 외 arxiv

Physical Reservoir Computing (PRC) leverages the intrinsic nonlinear dynamics of physical substrates, mechanical, optical, spintronic, and beyond, as fixed computational reservoirs, offering a compelling paradigm for ene…

An Extensible Benchmark Suite for Learning to Simulate Physical Systems

2021-08-09 · Karl Otness, Arvi Gjoka, Joan Bruna, Daniele Panozzo 외

Simulating physical systems is a core component of scientific computing, encompassing a wide range of physical domains and applications. Recently, there has been a surge in data-driven methods to complement traditional n…

Computational EfficiencyDiversity

NEWTON: Are Large Language Models Capable of Physical Reasoning?

2023-10-10 · Yi Ru Wang, Jiafei Duan, Dieter Fox, Siddhartha Srinivasa

Large Language Models (LLMs), through their contextualized representations, have been empirically proven to encapsulate syntactic, semantic, word sense, and common-sense knowledge. However, there has been limited explora…

AttributeCommon Sense Reasoning