College Mathematics
1개 벤치마크 · 논문 4편 · 이 태스크의 논문 보기 →
Benchmarks
BIG-bench
Most implemented
Scaling Language Models: Methods, Analysis & Insights from Training Gopher
Papers
Evaluating AI Grading on Real-World Handwritten College Mathematics: A Large-Scale Study Toward a Benchmark
Grading in large undergraduate STEM courses often yields minimal feedback due to heavy instructional workloads. We present a large-scale empirical study of AI grading on real, handwritten single-variable calculus work fr…
Mathematical ReasoningCollege MathematicsEffectiveness of Zero-shot-CoT in Japanese Prompts
We compare the effectiveness of zero-shot Chain-of-Thought (CoT) prompting in Japanese and English using ChatGPT-3.5 and 4o-mini. The technique of zero-shot CoT, which involves appending a phrase such as "Let's think ste…
Abstract AlgebraCollege MathematicsMMLUMulti-task Language UnderstandingMathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark
Recent advancements in large language models (LLMs) have showcased significant improvements in mathematics. However, traditional math benchmarks like GSM8k offer a unidimensional perspective, falling short in providing a…
College MathematicsGSM8KMathScaling Language Models: Methods, Analysis & Insights from Training Gopher
Language modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world. In this paper, we present an analysis o…
Abstract AlgebraAnachronismsAnalogical SimilarityAnalytic Entailment+143