paper-with-me

홈 › Papers

Finding the Sweet Spot: Trading Quality, Cost, and Speed During Inference-Time LLM Reflection

2025-10-23 · Jack Butler, Nikita Kozodoi, Zainab Afolabi, Brian Tyacke, Gaiar Baimuratov arxiv

As Large Language Models (LLMs) continue to evolve, practitioners face increasing options for enhancing inference-time performance without model retraining, including budget tuning and multi-step techniques like self-reflection. While these methods improve output quality, they create complex trade-offs among accuracy, cost, and latency that remain poorly understood across different domains. This paper systematically compares self-reflection and budget tuning across mathematical reasoning and translation tasks. We evaluate prominent LLMs, including Anthropic Claude, Amazon Nova, and Mistral families, along with other models under varying reflection depths and compute budgets to derive Pareto optimal performance frontiers. Our analysis reveals substantial domain dependent variation in self-reflection effectiveness, with performance gains up to 220\% in mathematical reasoning. We further investigate how reflection round depth and feedback mechanism quality influence performance across model families. To validate our findings in a real-world setting, we deploy a self-reflection enhanced marketing content localisation system at Lounge by Zalando, where it shows market-dependent effectiveness, reinforcing the importance of domain specific evaluation when deploying these techniques. Our results provide actionable guidance for selecting optimal inference strategies given specific domains and resource constraints. We open source our self-reflection implementation for reproducibility at https://github.com/aws-samples/sample-genai-reflection-for-bedrock.

📄 PDF Abstract BibTeX arXiv:2510.20653

Code (0)

등록된 구현이 없습니다.

Tasks

Mathematical Reasoning

Similar Papers 제목 키워드 기반

Towards Maximizing a Perceptual Sweet Spot

2022-01-05 · Pedro Izquierdo Lehmann, Rodrigo F. Cadiz, Carlos A. Sing Long

The sweet spot can be interpreted as the region where acoustic sources create a spatial auditory illusion. We study the problem of maximizing this sweet spot when reproducing a desired sound wave using an array of loudsp…

Finding the SWEET Spot: Analysis and Improvement of Adaptive Inference in Low Resource Settings

2023-06-04 · Daniel Rotem, Michael Hassid, Jonathan Mamou, Roy Schwartz

Adaptive inference is a simple method for reducing inference costs. The method works by maintaining multiple classifiers of different capacities, and allocating resources to each test instance according to its difficulty…

Prediction of reservoir key parameters in ‘sweet spot’ on the basis of particle swarm optimization to TCN-LSTM network

2022-11-05 · 2022 2022 11 · Fengcai Huo

In oil reservoirs, the sweet spot is found that the well could be positioned quickly and accurately, the drilling rate and the oil-gas production are increased, development cost is reduced. Among them, sorting, granulari…

Prediction

Controlling the Quality of Distillation in Response-Based Network Compression

2021-12-19 · Vibhas Vats, David Crandall

The performance of a distillation-based compressed network is governed by the quality of distillation. The reason for the suboptimal distillation of a large network (teacher) to a smaller network (student) is largely att…

Knowledge Distillation

Morphology Informed Selections for Subword Vocabulary Size

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Currently, guidance around selection of an optimal or appropriate subword vocabulary size is incomplete and confusing at best. Using a measure of subword-morpheme overlap, our analysis shows that one can find a "sweet sp…