paper-with-me

Papers

$α$-divergence Improves the Entropy Production Estimation via Machine Learning

2023-03-06 · Euijoon Kwon, Yongjoo Baek

Recent years have seen a surge of interest in the algorithmic estimation of stochastic entropy production (EP) from trajectory data via machine learning. A crucial element of such algorithms is the identification of a loss function whose minimization guarantees the accurate EP estimation. In this study, we show that there exists a host of loss functions, namely those implementing a variational representation of the $\alpha$-divergence, which can be used for the EP estimation. By fixing $\alpha$ to a value between $-1$ and $0$, the $\alpha$-NEEP (Neural Estimator for Entropy Production) exhibits a much more robust performance against strong nonequilibrium driving or slow dynamics, which adversely affects the existing method based on the Kullback-Leibler divergence ($\alpha = 0$). In particular, the choice of $\alpha = -0.5$ tends to yield the optimal results. To corroborate our findings, we present an exactly solvable simplification of the EP estimation problem, whose loss function landscape and stochastic properties give deeper intuition into the robustness of the $\alpha$-NEEP.

📄 PDF Abstract BibTeX arXiv:2303.02901

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Entropy Production in Machine Learning Under Fokker-Planck Probability Flow

2026-01-02 · Lennon Shikhman arxiv

Machine learning models deployed in nonstationary environments inevitably experience performance degradation due to data drift. While numerous drift detection heuristics exist, most lack a dynamical interpretation and pr…

Learning Stochastic Thermodynamics Directly from Correlation and Trajectory-Fluctuation Currents

2025-04-26 · Jinghao Lyu, Kyle J. Ray, James P. Crutchfield

Markedly increased computational power and data acquisition have led to growing interest in data-driven inverse dynamics problems. These seek to answer a fundamental question: What can we learn from time series measureme…

Semantic Faithfulness and Entropy Production Measures to Tame Your LLM Demons and Manage Hallucinations

2025-12-04 · Igor Halperin arxiv

Evaluating faithfulness of Large Language Models (LLMs) to a given task is a complex challenge. We propose two new unsupervised metrics for faithfulness evaluation using insights from information theory and thermodynamic…

Answer Generation

SAFE: Stable Alignment Finetuning with Entropy-Aware Predictive Control for Reinforcement Learning from Human Feedback (RLHF)

2026-02-04 · Dipan Maity arxiv

Proximal Policy Optimization (PPO) has been positioned by recent literature as the canonical method for the RL part of Reinforcement Learning from Human Feedback (RLHF). PPO performs well empirically but has a heuristic …

Reinforcement Learning

Direct Debiased Machine Learning via Bregman Divergence Minimization

2025-10-27 · Masahiro Kato arxiv

We develop a direct debiased machine learning framework comprising Neyman targeted estimation and generalized Riesz regression. Our framework unifies Riesz regression for automatic debiased machine learning, covariate ba…