paper-with-me

홈 › Papers

MorFiC: Fixing Value Miscalibration for Zero-Shot Quadruped Transfer

2026-03-15 · Prakhar Mishra, Amir Hossain Raj, Xuesu Xiao, Dinesh Manocha arxiv

Generalizing learned locomotion policies across quadrupedal robots with different morphologies remains a challenge. Policies trained on a single robot often fail when deployed on embodiments with different mass distributions, kinematics, joint limits, or actuation constraints, forcing per-robot retraining. Prior works have approached this primarily either by scaling via training across real or generated embodiments or using large architectures producing transferable policies which are both storage and compute heavy. We argue that both of these approaches circumvent a key failure mode in actor-critic learning: a shared value function tends to average incompatible value targets across embodiments, yielding miscalibrated advantages and fixing that helps transfer and reduce the compute cost. We present MorFiC, a reinforcement learning approach for zero-shot cross-morphology locomotion which fixes this issue using multiplicative critic conditioning on a morphology latent. Trained with a single source robot with morphology randomization in simulation, MorFiC achieves zero-shot transfer to total seven robots and reaching forward velocity competitive or surpassing scaling-based and morphology-conditioned PPO baselines for examples 1.98 m/s on AlienGo, where as additive PPO baselines remain below 0.65 m/s. MorFiC achieves this using single training robot within 2 hours of training time compared to tens or hundreds of hours required for scaling baselines training on multiple embodiments. We diagnose critic calibration via three metric: Explained variance, advantage sign-flip rate and policy gradient cosine, confirming our critic conditioning as determinant for policy update direction. Finally, we demonstrate zero-shot deployment on Unitree Go1 and Go2 robots without fine-tuning.

📄 PDF Abstract BibTeX arXiv:2603.14554

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Calibrate to Discriminate: Improve In-Context Learning with Label-Free Comparative Inference

2024-10-03 · Wei Cheng, Tianlu Wang, Yanmin Ji, Fan Yang 외

While in-context learning with large language models (LLMs) has shown impressive performance, we have discovered a unique miscalibration behavior where both correct and incorrect predictions are assigned the same level o…

In-Context Learning

Robust Calibration of Large Vision-Language Adapters

2024-07-18 · Balamurali Murugesan, Julio Silva-Rodriguez, Ismail Ben Ayed, Jose Dolz

This paper addresses the critical issue of miscalibration in CLIP-based model adaptation, particularly in the challenging scenario of out-of-distribution (OOD) samples, which has been overlooked in the existing literatur…

Prompt LearningTest-time Adaptation

Evaluating the Effectiveness of LLMs in Fixing Maintainability Issues in Real-World Projects

2025-02-04 · Henrique Nunes, Eduardo Figueiredo, Larissa Rocha, Sarah Nadi 외

Large Language Models (LLMs) have gained attention for addressing coding problems, but their effectiveness in fixing code maintainability remains unclear. This study evaluates LLMs capability to resolve 127 maintainabili…

On the Calibration of Massively Multilingual Language Models

2022-10-21 · Kabir Ahuja, Sunayana Sitaram, Sandipan Dandapat, Monojit Choudhury

Massively Multilingual Language Models (MMLMs) have recently gained popularity due to their surprising effectiveness in cross-lingual transfer. While there has been much work in evaluating these models for their performa…

Cross-Lingual Transfer

Enabling Calibration In The Zero-Shot Inference of Large Vision-Language Models

2023-03-11 · Will LeVine, Benjamin Pikus, Pranav Raja, Fernando Amat Gil

Calibration of deep learning models is crucial to their trustworthiness and safe usage, and as such, has been extensively studied in supervised classification models, with methods crafted to decrease miscalibration. Howe…