paper-with-me

홈 › Papers

Iterative LLM-Based Generation and Refinement of Distracting Conditions in Math Word Problems

2025-10-08 · Kaiqi Yang, Hang Li, Yucheng Chu, Zitao Liu, Mi Tian, Hui Liu arxiv

Mathematical reasoning serves as a crucial testbed for the intelligence of large language models (LLMs), and math word problems (MWPs) are a popular type of math problems. Most MWP datasets consist of problems containing only the necessary information, while problems with distracting and excessive conditions are often overlooked. Prior works have tested popular LLMs and found a dramatic performance drop in the presence of distracting conditions. However, datasets of MWPs with distracting conditions are limited, and most suffer from lower levels of difficulty and out-of-context expressions. This makes distracting conditions easy to identify and exclude, thus reducing the credibility of benchmarking on them. Moreover, when adding distracting conditions, the reasoning and answers may also change, requiring intensive labor to check and write the solutions. To address these issues, we design an iterative framework to generate distracting conditions using LLMs. We develop a set of prompts to revise MWPs from different perspectives and cognitive levels, encouraging the generation of distracting conditions as well as suggestions for further revision. Another advantage is the shared solutions between original and revised problems: we explicitly guide the LLMs to generate distracting conditions that do not alter the original solutions, thus avoiding the need to generate new solutions. This framework is efficient and easy to deploy, reducing the overhead of generating MWPs with distracting conditions while maintaining data quality.

📄 PDF Abstract BibTeX arXiv:2510.08615

Code (0)

등록된 구현이 없습니다.

Tasks

Mathematical Reasoning

Similar Papers 제목 키워드 기반

Spontaneous Reward Hacking in Iterative Self-Refinement

2024-07-05 · Jane Pan, He He, Samuel R. Bowman, Shi Feng

Language models are capable of iteratively improving their outputs based on natural language feedback, thus enabling in-context optimization of user preference. In place of human users, a second language model can be use…

Language ModelingLanguage Modelling

GryphOne: Symbol-Aware Masked Diffusion for Structural Refinement in Offline Handwritten Mathematical Expression Recognition

2026-02-03 · Takaya Kawakatsu, Ryo Ishiyama arxiv

Handwritten mathematical expression recognition (HMER) requires reasoning over diverse symbols and structures, yet autoregressive models struggle with exposure bias and syntax inconsistency. We present GryphOne, a discre…

Self-Refine: Iterative Refinement with Self-Feedback

2023-03-30 · NeurIPS 2023 11 · Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 외

Like humans, large language models (LLMs) do not always generate the best output on their first try. Motivated by how humans refine their written text, we introduce Self-Refine, an approach for improving initial outputs …

Mathematical ReasoningResponse Generation

MultiVis-Agent: A Multi-Agent Framework with Logic Rules for Reliable and Comprehensive Cross-Modal Data Visualization

2026-01-26 · Jinwei Lu, Yuanfeng Song, Chen Zhang, Raymond Chi-Wing Wong arxiv

Real-world visualization tasks involve complex, multi-modal requirements that extend beyond simple text-to-chart generation, requiring reference images, code examples, and iterative refinement. Current systems exhibit fu…

The Distracting Effect: Understanding Irrelevant Passages in RAG

2025-05-11 · Chen Amiraz, Florin Cuconasu, Simone Filice, Zohar Karnin

A well-known issue with Retrieval Augmented Generation (RAG) is that retrieved passages that are irrelevant to the query sometimes distract the answer-generating LLM, causing it to provide an incorrect response. In this …

Binary ClassificationRAGRetrieval-augmented Generation