paper-with-me

GSM8K

1개 벤치마크 · 논문 439편 · 이 태스크의 논문 보기 →

Benchmarks

GSM8K

결과 29개

Most implemented

Qwen2 Technical Report

2024-07-15 · 구현 6개

Large Language Models as Optimizers

2023-09-07 · 구현 4개

Papers

GEMMAS: Graph-based Evaluation Metrics for Multi Agent Systems

2025-07-17 · Jisoo Lee, Raeyoung Chang, Dongwook Kwon, Harmanpreet Singh 외

Multi-agent systems built on language models have shown strong performance on collaborative reasoning tasks. However, existing evaluations focus only on the correctness of the final output, overlooking how inefficient co…

DiversityGSM8K

DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression

2025-07-16 · Yi Zhao, Zuchao Li, Hai Zhao, Baoyuan Qi 외

Task-agnostic prompt compression leverages the redundancy in natural language to reduce computational overhead and enhance information density within prompts, especially in long-context scenarios. Existing methods predom…

GSM8K

KisMATH: Do LLMs Have Knowledge of Implicit Structures in Mathematical Reasoning?

2025-07-15 · Soumadeep Saha, Akshay Chaturvedi, Saptarshi Saha, Utpal Garain 외

Chain-of-thought traces have been shown to improve performance of large language models in a plethora of reasoning tasks, yet there is no consensus on the mechanism through which this performance boost is achieved. To sh…

GSM8KLanguage ModelingLanguage ModellingMathematical Reasoning

CoRE: Enhancing Metacognition with Label-free Self-evaluation in LRMs

2025-07-08 · Haoxi Li, Sikai Bai, Jie Zhang, Song Guo

Large reasoning models (LRMs) have demonstrated impressive capabilities in domains like mathematics and program synthesis. Despite their strong performance, LRMs often exhibit overthinking -- excessive and redundant reas…

GSM8KMathMathematical ReasoningProgram Synthesis

any4: Learned 4-bit Numeric Representation for LLMs

2025-07-07 · Mostafa Elhoushi, Jeff Johnson

We present any4, a learned 4-bit weight quantization solution for large language models (LLMs) providing arbitrary numeric representations without requiring pre-processing of weights or activations. any4 yields higher ac…

GPUGSM8KHumanEvalmbpp+2

Activation Steering for Chain-of-Thought Compression

2025-07-07 · Seyedarmin Azizi, Erfan Baghaei Potraghloo, Massoud Pedram

Large language models (LLMs) excel at complex reasoning when they include intermediate steps, known as "chains of thought" (CoTs). However, these rationales are often overly verbose, even for simple problems, leading to …

GSM8KMath

전체 439편 보기 →