paper-with-me

홈 › Papers

CReF: Cross-modal and Recurrent Fusion for Depth-conditioned Humanoid Locomotion

2026-03-31 · Yuan Hao, Ruiqi Yu, Shixin Luo, Guoteng Zhang, Jun Wu, Qiuguo Zhu arxiv

Stable traversal over geometrically complex terrain increasingly requires exteroceptive perception, yet prior perceptive humanoid locomotion methods often remain tied to explicit geometric abstractions, either by mediating control through robot-centric 2.5D terrain representations or by shaping depth learning with auxiliary geometry-related targets. Such designs inherit the representational bias of the intermediate or supervisory target and can be restrictive for vertical structures, perforated obstacles, and complex real-world clutter. We propose CReF (Cross-modal and Recurrent Fusion), a single-stage depth-conditioned humanoid locomotion framework that learns locomotion-relevant features directly from raw forward-facing depth without explicit geometric intermediates. CReF couples proprioception and depth tokens through proprioception-queried cross-modal attention, fuses the resulting representation with a gated residual fusion block, and performs temporal integration with a Gated Recurrent Unit (GRU) regulated by a highway-style output gate for state-dependent blending of recurrent and feedforward features. To further improve terrain interaction, we introduce a terrain-aware foothold placement reward that extracts supportable foothold candidates from foot-end point-cloud samples and rewards touchdown locations that lie close to the nearest supportable candidate. Experiments in simulation and on a physical humanoid demonstrate robust traversal over diverse terrains and effective zero-shot transfer to real-world scenes containing handrails, hollow pallet assemblies, severe reflective interference, and visually cluttered outdoor surroundings.

📄 PDF Abstract BibTeX arXiv:2603.29452

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DocRefine: An Intelligent Framework for Scientific Document Understanding and Content Optimization based on Multimodal Large Model Agents

2025-08-09 · Kun Qian, Wenjie Li, Tianyu Sun, Wenhong Wang 외 arxiv

The exponential growth of scientific literature in PDF format necessitates advanced tools for efficient and accurate document understanding, summarization, and content optimization. Traditional methods fall short in hand…

LocRef-Diffusion:Tuning-Free Layout and Appearance-Guided Generation

2024-11-22 · Fan Deng, Yaguang Wu, Xinyang Yu, Xiangjun Huang 외

Recently, text-to-image models based on diffusion have achieved remarkable success in generating high-quality images. However, the challenge of personalized, controllable generation of instances within these images remai…

IncreFA: Breaking the Static Wall of Generative Model Attribution

2026-04-20 · Haotian Qin, Dongliang Chang, Yueying Gao, Yuexuan Tan 외 arxiv

As AI generative models evolve at unprecedented speed, image attribution has become a moving target. New diffusion, adversarial and autoregressive generators appear almost monthly, making existing watermark, classifier a…

Incremental LearningImage Attribution

SpecRef: A Fast Training-free Baseline of Specific Reference-Condition Real Image Editing

2024-01-07 · Songyan Chen, Jiancheng Huang

Text-conditional image editing based on large diffusion generative model has attracted the attention of both the industry and the research community. Most existing methods are non-reference editing, with the user only ab…

Multimodel-guided image editingText-based Image Editingtext-guided-image-editingText-to-Image Generation

Speculative Refinement: A Hybrid Autoregressive Diffusion Decoding Strategy and Its Behavior Across Benchmarks

2026-06-25 · Aditi Gupta, Neel Mishra, Kushagra Trivedi, Pawan Kumar arxiv

How should we evaluate generation systems that combine autoregressive (AR) and diffusion decoding? We study this question through Speculative Refinement (SpecRef), a training-free hybrid method that warm-starts a masked …