paper-with-me

홈 › Papers

TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research Ability

2026-08-10 · Vincent Cohen-Addad, Dimitris Paparas, Ernest van Wijland, Max Springer, Julien Canitrot-Paradis, Honghao Lin, David Woodruff, Adarsh Kumarappan, Rajesh Jayaram, Rudrajit Das, Lalit Jain, Ola Svensson, Silvio Lattanzi, Mislav Balunovic, Theophane Weber, Vahab Mirrokni arxiv

We introduce TCS-Bench, a benchmark for evaluating Large Language Models (LLMs) on research-level Theoretical Computer Science (TCS) proof generation. TCS-Bench consists of theorem-proving tasks from papers published at top theoretical computer science venues (STOC, FOCS, and SODA). Each task provides the necessary context to derive a self-contained proof for a target result. We evaluate state-of-the-art models on this benchmark. We verify the correctness of generated proofs via a verification agent, and further benchmark the verifier against human-expert proof judgements on a set of target statements and generated proofs pairs. Our reference verifier achieves over 90% accuracy on the expert labeled set.

📄 PDF Abstract BibTeX arXiv:2608.09538

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Learnability and Complexity of Quantum Samples

2020-10-22 · Murphy Yuezhen Niu, Andrew M. Dai, Li Li, Augustus Odena 외

Given a quantum circuit, a quantum computer can sample the output distribution exponentially faster in the number of bits than classical computers. A similar exponential separation has yet to be established in generative…

Benchmarking

Cloud to Edge: Benchmarking LLM Inference On Hardware-Accelerated Single-Board Computers

2026-04-24 · Harri Renney, Fouad Trad, Michael Mattarock, Zena Wood arxiv

Large language models (LLMs) are becoming increasingly capable at small parameter scales. At the same time, conventional cloud-centric deployment introduces challenges around data privacy, latency, and cost that are acut…

The Benchmarking Epistemology: Construct Validity for Evaluating Machine Learning Models

2025-10-27 · Timo Freiesleben, Sebastian Zezulka arxiv

Predictive benchmarking, the evaluation of machine learning models based on predictive performance and competitive ranking, is a central epistemic practice in machine learning research and an increasingly prominent metho…

Image Classification

Featuremetric benchmarking: Quantum computer benchmarks based on circuit features

2025-04-17 · Timothy Proctor, Anh Tran, Xingxin Liu, Aditya Dhumuntarao 외

Benchmarks that concisely summarize the performance of many-qubit quantum computers are essential for measuring progress towards the goal of useful quantum computation. In this work, we present a benchmarking framework t…

Benchmarking

OSWorld-Human: Benchmarking the Efficiency of Computer-Use Agents

2025-06-19 · Reyna Abhyankar, Qi Qi, Yiying Zhang

Generative AI is being leveraged to solve a variety of computer-use tasks involving desktop applications. State-of-the-art systems have focused solely on improving accuracy on leading benchmarks. However, these systems a…

Benchmarking