GSM8K
1개 벤치마크 · 논문 439편 · 이 태스크의 논문 보기 →
Benchmarks
GSM8K
Most implemented
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
Qwen2 Technical Report
Training Verifiers to Solve Math Word Problems
Large Language Models as Optimizers
Language Models are Multilingual Chain-of-Thought Reasoners
Papers
GEMMAS: Graph-based Evaluation Metrics for Multi Agent Systems
Multi-agent systems built on language models have shown strong performance on collaborative reasoning tasks. However, existing evaluations focus only on the correctness of the final output, overlooking how inefficient co…
DiversityGSM8KDAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression
Task-agnostic prompt compression leverages the redundancy in natural language to reduce computational overhead and enhance information density within prompts, especially in long-context scenarios. Existing methods predom…
GSM8KKisMATH: Do LLMs Have Knowledge of Implicit Structures in Mathematical Reasoning?
Chain-of-thought traces have been shown to improve performance of large language models in a plethora of reasoning tasks, yet there is no consensus on the mechanism through which this performance boost is achieved. To sh…
GSM8KLanguage ModelingLanguage ModellingMathematical ReasoningCoRE: Enhancing Metacognition with Label-free Self-evaluation in LRMs
Large reasoning models (LRMs) have demonstrated impressive capabilities in domains like mathematics and program synthesis. Despite their strong performance, LRMs often exhibit overthinking -- excessive and redundant reas…
GSM8KMathMathematical ReasoningProgram Synthesisany4: Learned 4-bit Numeric Representation for LLMs
We present any4, a learned 4-bit weight quantization solution for large language models (LLMs) providing arbitrary numeric representations without requiring pre-processing of weights or activations. any4 yields higher ac…
GPUGSM8KHumanEvalmbpp+2Activation Steering for Chain-of-Thought Compression
Large language models (LLMs) excel at complex reasoning when they include intermediate steps, known as "chains of thought" (CoTs). However, these rationales are often overly verbose, even for simple problems, leading to …
GSM8KMath