paper-with-me

Papers

Athena-WBC: Capability-Aligned Policy Experts for Long-Tail Humanoid Whole-Body Control

2026-07-06 · Yuan Jiang, Ningyuan Zhang, Xicun Yang, Shidi Li, Yuzhi Jiang, Zhiyi Rong, Shuaikang Ma, Chuanzheng Li, Jie Chen arxiv

Large-scale humanoid motion-tracking controllers are commonly improved by reallocating training effort: difficult motions are sampled more often, isolated into smaller subsets, or assigned to specialized experts. We show that this view is incomplete. In strong whole-body-control baselines, a residual set of feasible training clips remains unsolved even under targeted training, especially for high-dynamic transitions and balance-critical motions. These failures arise not only from insufficient exposure, but from a mismatch between the motion demands and the effective capability induced by the default training recipe. We propose Athena-WBC, a compact teacher-student pipeline with capability-aligned policy experts for long-tail humanoid whole-body control. Dynamic experts use a tracking-focused, constraint-aware objective that removes conservative effort and temporal-control penalties while preserving physical feasibility constraints; balance experts use a gravity curriculum to improve early-training survivability. The resulting privileged teachers are motion-routed for DAgger distillation and then compressed into a single controller with deployable observations followed by RL fine-tuning. Experiments on a full-size humanoid show improved recovery of training-set long-tail motions and better held-out tracking than a strong SONIC-recipe baseline, using only a small number of experts.

📄 PDF Abstract BibTeX arXiv:2607.04837

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Athena: Synergizing Data Prefetching and Off-Chip Prediction via Online Reinforcement Learning

2026-01-24 · Rahul Bera, Zhenrong Lang, Caroline Hengartner, Konstantinos Kanellopoulos 외 arxiv

Prefetching and off-chip prediction are two techniques proposed to hide long memory access latencies in high-performance processors. In this work, we demonstrate that: (1) prefetching and off-chip prediction often provid…

Reinforcement Learning

A Transformer-based Response Evaluator for Open-Domain Spoken Conversation

2023-02-09 · Vrindavan Harrison, Rishi Rajasekaran, Marilyn Walker

Many open-domain dialogue systems rely on multiple response generators, any of which can contribute a response to the dialogue in a particular context. Thus the ability to compare potential responses and then select the …

What Would an LLM Do? Evaluating Large Language Models for Policymaking to Alleviate Homelessness

2025-09-04 · Pierre Le Coz, Jia An Liu, Debarun Bhattacharjya, Georgina Curto 외 arxiv

Large language models (LLMs) are increasingly being adopted in high-stakes domains. Their potential to encode evolving social contexts and to generate plausible scenarios position them as promising tools in social policy…

ATHENA: Agentic Team for Hierarchical Evolutionary Numerical Algorithms

2025-12-03 · Juan Diego Toscano, Daniel T. Chen, George Em Karniadakis arxiv

Progress in computational science depends on complex numerical workflows that must faithfully encode physical laws, yet translating conceptual insight into reliable code remains a major bottleneck. Although large languag…

Athena 2.0: Contextualized Dialogue Management for an Alexa Prize SocialBot

2021-11-03 · EMNLP (ACL) 2021 11 · Juraj Juraska, Kevin K. Bowden, Lena Reed, Vrindavan Harrison 외

Athena 2.0 is an Alexa Prize SocialBot that has been a finalist in the last two Alexa Prize Grand Challenges. One reason for Athena's success is its novel dialogue management strategy, which allows it to dynamically cons…

Dialogue ManagementManagement