paper-with-me

College Mathematics

1개 벤치마크 · 논문 4편 · 이 태스크의 논문 보기 →

Benchmarks

BIG-bench

결과 1개

Most implemented

Papers

Evaluating AI Grading on Real-World Handwritten College Mathematics: A Large-Scale Study Toward a Benchmark

2026-03-01 · Zhiqi Yu, Xingping Liu, Haobin Mao, Mingshuo Liu 외 arxiv

Grading in large undergraduate STEM courses often yields minimal feedback due to heavy instructional workloads. We present a large-scale empirical study of AI grading on real, handwritten single-variable calculus work fr…

Mathematical ReasoningCollege Mathematics

Effectiveness of Zero-shot-CoT in Japanese Prompts

2025-03-09 · Shusuke Takayama, Ian Frank

We compare the effectiveness of zero-shot Chain-of-Thought (CoT) prompting in Japanese and English using ChatGPT-3.5 and 4o-mini. The technique of zero-shot CoT, which involves appending a phrase such as "Let's think ste…

Abstract AlgebraCollege MathematicsMMLUMulti-task Language Understanding

MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

2024-05-20 · Hongwei Liu, Zilong Zheng, Yuxuan Qiao, Haodong Duan 외

Recent advancements in large language models (LLMs) have showcased significant improvements in mathematics. However, traditional math benchmarks like GSM8k offer a unidimensional perspective, falling short in providing a…

College MathematicsGSM8KMath

Scaling Language Models: Methods, Analysis & Insights from Training Gopher

2021-12-08 · NA 2021 12 · Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican 외

Language modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world. In this paper, we present an analysis o…

Abstract AlgebraAnachronismsAnalogical SimilarityAnalytic Entailment+143