paper-with-me

홈 › Papers

Where Paths Split: Localized, Calibrated Control of Moral Reasoning in Large Language Models

2026-05-05 · Chenchen Yuan, Zheyu Zhang, Gjergji Kasneci arxiv

Large language models often display heterogeneous moral preferences across settings. We study inference-time steering toward a desired ethical framework while preserving general competence. We present Convergent-Divergent Routing, which traces and edits minimal branch points inside transformer blocks where ethical-framework-related pathways first converge and then diverge. Gating non-target branches at these loci blocks the downstream propagation while leaving upstream computations intact. We find that this intervention alone increases targeted ethical-framework reasoning. To achieve fine-grained control, we adapt Common Spatial Patterns to the residual stream and extract, for each branch-point layer, a pair of directions that discriminate between utilitarian and deontological frameworks. We then introduce Dual Logit Calibration, a closed-form, minimum-$\ell_2$-norm update that moves the residual within this two-dimensional subspace so the resulting directional projections align with user-specified preference weights. Experiments on real-life moral dilemmas show that our method reliably achieves preference calibration and largely preserves general capabilities, outperforming recent baselines while providing an interpretable mechanism.

📄 PDF Abstract BibTeX arXiv:2605.03609

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Where Is My Physics Wrong? Localized and Identifiable Discovery of Model Discrepancy

2026-06-22 · Yifan Wang arxiv

Hybrid models combine trusted physics with data-driven correction, but a physical model is rarely wrong everywhere or in the same way. The key diagnostic question is local: where does the model fail, what missing mechani…

Predictive Memory Localization: Forecasting Selective Intervention Paths from Internal Signals

2026-08-13 · Jinhao Jing, Tian Zeyu, Lucas Qingyang Fang, Zhisheng Chen 외 arxiv

Activation steering turns localized representations into control directions, but localization alone does not reveal whether a direction has a selective operating regime. We introduce Predictive Memory Localization (PML),…

Double Descent in Gradient Boosting Decision Trees via Split-Candidate Scaling

2026-08-04 · Ryuichi Kanoh arxiv

Double descent is commonly studied by scaling an explicit capacity parameter, such as neural-network width. For gradient boosting decision trees (GBDTs), however, an analogous single-axis capacity parameter has not been …

Direct Learning of Calibration-Aware Uncertainty for Neural PDE Surrogates

2026-02-11 · Carlos Stein Brito arxiv

Neural PDE surrogates are often deployed in data-limited or partially observed regimes where downstream decisions depend on calibrated uncertainty in addition to low prediction error. Existing approaches obtain uncertain…

Localized Adaptive Risk Control

2024-05-13 · Matteo Zecchin, Osvaldo Simeone

Adaptive Risk Control (ARC) is an online calibration strategy based on set prediction that offers worst-case deterministic long-term risk control, as well as statistical marginal coverage guarantees. ARC adjusts the size…

ARCFairnessImage SegmentationPrediction+1