paper-with-me

Papers

SPAR: Support-Preserving Action Rectification

2026-05-27 · Jiaxin Zhao, Weihang Pan, Xun Liang, Binbin Lin arxiv

Offline policy improvement faces an inherent conflict between maximizing value and fitting the data distribution. While in-sample weighted regression is stable, it suffers from over-conservatism that suppresses high-value actions in the distribution tail; conversely, gradient-based approaches often exhibit a fitting-optimization conflict of gradients, which drives the policy off the data manifold. To address this, we propose Support-Preserving Action Rectification (SPAR), which reframes global learning as a local residual rectification anchored to a frozen pure behavior cloning policy. This framework performs fine-grained fitting and local policy improvement in the residual space, thereby contracting the search space. We further introduce Latent Self-Imitation, utilizing a latent-sampling weighted-regression mechanism to address fitting-improvement gradient conflict in the residual space. Theoretically, we prove this mechanism eliminates the manifold-normal drift of standard value gradients, while extensive D4RL experiments show SPAR extracts significant gains from suboptimal baselines to achieve state-of-the-art performance.

📄 PDF Abstract BibTeX arXiv:2605.27877

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Rectified LpJEPA: Joint-Embedding Predictive Architectures with Sparse and Maximum-Entropy Representations

2026-02-01 · Yilun Kuang, Yash Dagade, Tim G. J. Rudner, Randall Balestriero 외 arxiv

Joint-Embedding Predictive Architectures (JEPA) learn view-invariant representations and admit projection-based distribution matching for collapse prevention. Existing approaches regularize representations towards isotro…

Image Classification

GFT: From Imitation to Reward Fine-Tuning with Unbiased Group Advantages and Dynamic Coefficient Rectification

2026-04-15 · Wangjie Gan, Miao Pan, Linbo Xi, Wenqi Zhang 외 arxiv

Large language models are typically post-trained using supervised fine-tuning (SFT) and reinforcement learning (RL), yet effectively unifying efficient knowledge injection with robust generalization remains challenging. …

Reinforcement Learning

Turn Waste into Worth: Rectifying Top-$k$ Router of MoE

2024-02-17 · Zhiyuan Zeng, Qipeng Guo, Zhaoye Fei, Zhangyue Yin 외

Sparse Mixture of Experts (MoE) models are popular for training large language models due to their computational efficiency. However, the commonly used top-$k$ routing mechanism suffers from redundancy computation and me…

Computational EfficiencyGPUMixture-of-Experts

DREAM: Diffusion Rectification and Estimation-Adaptive Models

2023-11-30 · CVPR 2024 1 · Jinxin Zhou, Tianyu Ding, Tianyi Chen, Jiachen Jiang 외

We present DREAM, a novel training framework representing Diffusion Rectification and Estimation Adaptive Models, requiring minimal code changes (just three lines) yet significantly enhancing the alignment of training wi…

Image Super-ResolutionSuper-Resolution

ScenePilot: Grow-and-Repair Policy for Text-Driven 3D Indoor Scene Generation

2026-08-31 · Jiawei Zhang, Hongsong Wang, Pan Zhou arxiv

Text-driven 3D indoor scene generation has advanced from dataset-bound layout modeling to open-vocabulary synthesis with large language and vision-language models. Yet existing methods remain limited: one-pass generators…

Scene Generation