paper-with-me

Papers Mathematical Problem-Solving

“Mathematical Problem-Solving” 태그가 달린 논문 106편 · 필터 해제

EvoAgentX: An Automated Framework for Evolving Agentic Workflows

2025-07-04 · Yingxu Wang, Siwei Liu, Jinyuan Fang, Zaiqiao Meng

Multi-agent systems (MAS) have emerged as a powerful paradigm for orchestrating large language models (LLMs) and specialized tools to collaboratively address complex tasks. However, existing MAS frameworks often require …

Code GenerationMathMathematical Problem-Solvingmbpp

LocationReasoner: Evaluating LLMs on Real-World Site Selection Reasoning

2025-06-16 · Miho Koda, Yu Zheng, Ruixian Ma, Mingyang Sun 외

Recent advances in large language models (LLMs), particularly those enhanced through reinforced post-training, have demonstrated impressive reasoning capabilities, as exemplified by models such as OpenAI o1 and DeepSeek-…

Code GenerationMathematical Problem-Solving

TeleMath: A Benchmark for Large Language Models in Telecom Mathematical Problem Solving

2025-06-12 · Vincenzo Colle, Mohamed Sana, Nicola Piovesan, Antonio De Domenico 외

The increasing adoption of artificial intelligence in telecommunications has raised interest in the capability of Large Language Models (LLMs) to address domain-specific, mathematically intensive tasks. Although recent a…

Logical ReasoningMathematical Problem-SolvingMathematical Reasoning

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning

2025-06-10 · Xiao Liang, Zhong-Zhi Li, Yeyun Gong, Yang Wang 외

Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective for training large language models (LLMs) on complex reasoning tasks, such as mathematical problem solving. A prerequisite for the scalability of…

Knowledge DistillationMathMathematical Problem-Solving

Solving Inequality Proofs with Large Language Models

2025-06-09 · Jiayi Sheng, Luna Lyu, Jikai Jin, Tony Xia 외

Inequality proving, crucial across diverse scientific and mathematical fields, tests advanced reasoning skills such as discovering tight bounds and strategic theorem application. This makes it a distinct, demanding front…

Mathematical Problem-SolvingRelation Prediction

Chain-of-Code Collapse: Reasoning Failures in LLMs via Adversarial Prompting in Code Generation

2025-06-08 · Jaechul Roh, Varun Gandhi, Shivani Anilkumar, Arin Garg

Large Language Models (LLMs) have achieved remarkable success in tasks requiring complex reasoning, such as code generation, mathematical problem solving, and algorithmic synthesis -- especially when aided by reasoning t…

Code GenerationMathematical Problem-Solving

MORSE-500: A Programmatically Controllable Video Benchmark to Stress-Test Multimodal Reasoning

2025-06-05 · Zikui Cai, Andrew Wang, Anirudh Satheesh, Ankit Nakhawa 외

Despite rapid advances in vision-language models (VLMs), current benchmarks for multimodal reasoning fall short in three key dimensions. First, they overwhelmingly rely on static images, failing to capture the temporal c…

Dataset GenerationMathematical Problem-SolvingMultimodal Reasoning

PoLAR: Polar-Decomposed Low-Rank Adapter Representation

2025-06-03 · Kai Lion, Liang Zhang, Bingcong Li, Niao He

We show that low-rank adaptation of large-scale models suffers from a low stable rank that is well below the linear algebraic rank of the subspace, degrading fine-tuning performance. To mitigate the underutilization of t…

Mathematical Problem-SolvingRiemannian optimization

Evaluation of LLMs for mathematical problem solving

2025-05-30 · Ruonan Wang, Runxi Wang, Yunwen Shen, Chengfeng Wu 외

Large Language Models (LLMs) have shown impressive performance on a range of educational tasks, but are still understudied for their potential to solve mathematical problems. In this study, we compare three prominent LLM…

GSM8KMathematical Problem-SolvingMathematical Reasoning

Decomposing Elements of Problem Solving: What "Math" Does RL Teach?

2025-05-28 · Tian Qin, Core Francisco Park, Mujin Kwun, Aaron Walsman 외

Mathematical reasoning tasks have become prominent benchmarks for assessing the reasoning capabilities of LLMs, especially with reinforcement learning (RL) methods such as GRPO showing significant performance gains. Howe…

MathMathematical Problem-SolvingMathematical ReasoningReinforcement Learning (RL)

Surrogate Signals from Format and Length: Reinforcement Learning for Solving Mathematical Problems without Ground Truth Answers

2025-05-26 · Rihui Xin, Han Liu, Zecheng Wang, Yupeng Zhang 외

Large Language Models have achieved remarkable success in natural language processing tasks, with Reinforcement Learning playing a key role in adapting them to specific applications. However, obtaining ground truth answe…

Logical ReasoningMathematical Problem-Solving

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision

2025-05-26 · Tej Deep Pala, Panshul Sharma, Amir Zadeh, Chuan Li 외

Large Language Models (LLMs) are prone to hallucination, especially during multi-hop and reasoning-intensive tasks such as mathematical problem solving. While Outcome Reward Models verify only final answers, Process Rewa…

HallucinationMathMathematical Problem-SolvingMathematical Reasoning+1

RaDeR: Reasoning-aware Dense Retrieval Models

2025-05-23 · Debrup Das, Sam O' Nuallain, Razieh Rahimi

We propose RaDeR, a set of reasoning-based dense retrieval models trained with data derived from mathematical problem solving using large language models (LLMs). Our method leverages retrieval-augmented reasoning traject…

MathMathematical Problem-SolvingMathematical ReasoningRetrieval

SMART: Self-Generating and Self-Validating Multi-Dimensional Assessment for LLMs' Mathematical Problem Solving

2025-05-22 · Yujie Hou, Ting Zhang, Mei Wang, Xuetao Ma 외

Large Language Models have achieved remarkable results on a variety of mathematical benchmarks. However, concerns remain as to whether these successes reflect genuine mathematical reasoning or superficial pattern recogni…

DiagnosticMathematical Problem-SolvingMathematical Reasoning

Can reasoning models comprehend mathematical problems in Chinese ancient texts? An empirical study based on data from Suanjing Shishu

2025-05-22 · Liu Chang, Wang Dongbo, Liu Liu, Zhao Zhixiao

This study addresses the challenges in intelligent processing of Chinese ancient mathematical classics by constructing Guji_MATH, a benchmark for evaluating classical texts based on Suanjing Shishu. It systematically ass…

Mathematical Problem-Solving

Towards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems

2025-05-21 · Chengwei Wei, Bin Wang, Jung-jae Kim, Nancy F. Chen

Recent advances in large language models (LLMs) and multimodal LLMs (MLLMs) have led to strong reasoning ability across a wide range of tasks. However, their ability to perform mathematical reasoning from spoken input re…

BenchmarkingMathMathematical Problem-SolvingMathematical Reasoning+1

HARDMath2: A Benchmark for Applied Mathematics Built by Students as Part of a Graduate Class

2025-05-17 · James V. Roggeveen, Erik Y. Wang, Will Flintoft, Peter Donets 외

Large language models (LLMs) have shown remarkable progress in mathematical problem-solving, but evaluation has largely focused on problems that have exact analytical solutions or involve formal proofs, often overlooking…

MathMathematical Problem-SolvingMathematical Reasoning

Is PRM Necessary? Problem-Solving RL Implicitly Induces PRM Capability in LLMs

2025-05-16 · Zhangying Feng, Qianglong Chen, Ning Lu, YongQian Li 외

The development of reasoning capabilities represents a critical frontier in large language models (LLMs) research, where reinforcement learning (RL) and process reward models (PRMs) have emerged as predominant methodolog…

Mathematical Problem-SolvingReinforcement Learning (RL)

Reasoning with OmniThought: A Large CoT Dataset with Verbosity and Cognitive Difficulty Annotations

2025-05-16 · Wenrui Cai, Chengyu Wang, Junbing Yan, Jun Huang 외

The emergence of large reasoning models (LRMs) has transformed Natural Language Processing by excelling in complex tasks such as mathematical problem-solving and code generation. These models leverage chain-of-thought (C…

Code GenerationMathematical Problem-Solving

PT-MoE: An Efficient Finetuning Framework for Integrating Mixture-of-Experts into Prompt Tuning

2025-05-14 · Zongqian Li, Yixuan Su, Nigel Collier

Parameter-efficient fine-tuning (PEFT) methods have shown promise in adapting large language models, yet existing approaches exhibit counter-intuitive phenomena: integrating router into prompt tuning (PT) increases train…

MathMathematical Problem-SolvingMixture-of-Expertsparameter-efficient fine-tuning+1
1–20 / 106 다음 →