paper-with-me

홈 › Papers

DEL: Digit Entropy Loss for Numerical Learning of Large Language Models

2026-05-19 · Zhaohui Zheng, Chenhang He, Shihao Wang, Yuxuan Li, Ming-Ming Cheng, Lei Zhang arxiv

Number prediction stands as a fundamental capability of large language models (LLMs) in mathematical problem-solving and code generation. The widely adopted maximum likelihood estimation (MLE) for LLM training is not tailored to number prediction. Recently, penalty-driven approaches, e.g., Number Token Loss and Discretized Distance Loss, introduce an inductive bias of numerical distance but induce over-sharpened and over-flattened digit distributions, respectively. In this paper, we make an in-depth analysis on LLM numerical learning, and show that existing numerical learning methods conceptually follow a criterion-distance formulation, where the criterion term represents optimization pattern and the distance term instills geometric prior. Consequently, we present Digit Entropy Loss (DEL) for auto-regressive numerical learning, which reformulates the conventional unsupervised entropy optimization in three key designs: leveraging digit conditional probability and binary cross-entropy to guide the entropy optimization into a supervised manner; deprecating the distance term to bypass the issue of numerical distance; and generalizing the integer-based numerical learning to floating-point number optimization, enabling more accurate number prediction. Our DEL formulation can incorporate integers, decimals, and decimal points, expanding the learning objective from a single digit to the floating-point number domain. Experiments conducted on seven mathematical reasoning benchmarks with four representative LLMs, including CodeLlama, Mistral, DeepSeek, and Qwen-2.5, demonstrate that DEL consistently outperforms its counterparts in both overall prediction accuracy and numerical distance. Source codes are at https://github.com/PolyU-VCLab/DEL

📄 PDF Abstract BibTeX arXiv:2605.20369

Code (0)

등록된 구현이 없습니다.

Tasks

Mathematical ReasoningCode Generation

Similar Papers 제목 키워드 기반

Advancing Sequential Numerical Prediction in Autoregressive Models

2025-05-19 · Xiang Fei, Jinghui Lu, Qi Sun, Hao Feng 외

Autoregressive models have become the de facto choice for sequence generation tasks, but standard approaches treat digits as independent tokens and apply cross-entropy loss, overlooking the coherent structure of numerica…

Prediction

Cut Your Losses in Large-Vocabulary Language Models

2024-11-13 · Erik Wijmans, Brody Huval, Alexander Hertzberg, Vladlen Koltun 외

As language models grow ever larger, so do their vocabularies. This has shifted the memory footprint of LLMs during training disproportionately to one single layer: the cross-entropy in the loss computation. Cross-entrop…

Adaptively Solving the Local-Minimum Problem for Deep Neural Networks

2020-12-25 · Huachuan Wang, James Ting-Ho Lo

This paper aims to overcome a fundamental problem in the theory and application of deep neural networks (DNNs). We propose a method to solve the local minimum problem in training DNNs directly. Our method is based on the…

A machine-learning-assisted progressive digit-randomness screening framework for detecting non-random patterns in raw numerical research data

2026-06-05 · Zhuphua Cao arxiv

Raw numerical datasets remain less systematically examined in integrity screening than images, plagiarism, or summary-statistic inconsistencies. We developed the Fabrication-risk Digit Randomness Screening model (FDRS), …

Relative Entropy-Based Constant-Envelope Beamforming for Target Detection in Large-Scale MIMO Radar With Low-Resoultion ADCs

2023-01-19 · Ziyang Cheng, Linlong Wu, Bowen Wang, Julan Xie 외

Hybrid digital/analog architecture and low-resolution analog-to-digital/digital-to-analog converters (ADCs /DACs) are two low-cost implementations for large-scale millimeter wave (mmWave) systems. In this paper, we inves…