paper-with-me

홈 › Papers

TwinAligner: Visual-Dynamic Alignment Empowers Physics-aware Real2Sim2Real for Robotic Manipulation

2025-12-22 · Hongwei Fan, Hang Dai, Jiyao Zhang, Jinzhou Li, Qiyang Yan, Yujie Zhao, Mingju Gao, Jinghang Wu, Hao Tang, Hao Dong arxiv

The robotics field is evolving towards data-driven, end-to-end learning, inspired by multimodal large models. However, reliance on expensive real-world data limits progress. Simulators offer cost-effective alternatives, but the gap between simulation and reality challenges effective policy transfer. This paper introduces TwinAligner, a novel Real2Sim2Real system that addresses both visual and dynamic gaps. The visual alignment module achieves pixel-level alignment through SDF reconstruction and editable 3DGS rendering, while the dynamic alignment module ensures dynamic consistency by identifying rigid physics from robot-object interaction. TwinAligner improves robot learning by providing scalable data collection and establishing a trustworthy iterative cycle, accelerating algorithm development. Quantitative evaluations highlight TwinAligner's strong capabilities in visual and dynamic real-to-sim alignment. This system enables policies trained in simulation to achieve strong zero-shot generalization to the real world. The high consistency between real-world and simulated policy performance underscores TwinAligner's potential to advance scalable robot learning. Code and data will be released on https://twin-aligner.github.io

📄 PDF Abstract BibTeX arXiv:2512.19390

Code (0)

등록된 구현이 없습니다.

Tasks

Zero-shot Generalization

Similar Papers 제목 키워드 기반

Physics-Informed Video Generation via Mixture-of-Experts Latent Alignment

2026-06-03 · Cong Wang, Hanxin Zhu, Jiayi Luo, Yonglin Tian 외 arxiv

Large-scale video generation models have made remarkable progress in semantic consistency and visual quality, producing videos that are increasingly coherent and visually convincing. Nevertheless, the dynamics induced by…

Video Generation

Shrinking the Teacher: An Adaptive Teaching Paradigm for Asymmetric EEG-Vision Alignment

2025-11-14 · Lukun Wu, Jie Li, Ziqi Ren, Kaifan Zhang 외 arxiv

Decoding visual features from EEG signals is a central challenge in neuroscience, with cross-modal alignment as the dominant approach. We argue that the relationship between visual and brain modalities is fundamentally a…

Image Retrieval

UniVSE: Robust Visual Semantic Embeddings via Structured Semantic Representations

2019-04-11 · Hao Wu, Jiayuan Mao, Yufeng Zhang, Yuning Jiang 외

We propose Unified Visual-Semantic Embeddings (UniVSE) for learning a joint space of visual and textual concepts. The space unifies the concepts at different levels, including objects, attributes, relations, and full sce…

Contrastive LearningCross-Modal RetrievalRetrievalSentence

PhysAlign: Physics-Coherent Image-to-Video Generation through Feature and 3D Representation Alignment

2026-03-14 · Zhexiao Xiong, Yizhi Song, Liu He, Wei Xiong 외 arxiv

Video Diffusion Models (VDMs) offer a promising approach for simulating dynamic scenes and environments, with broad applications in robotics and media generation. However, existing models often generate temporally incohe…

Synthetic Data GenerationPhysical IntuitionVideo Generation

From One-to-One to Many-to-Many: Dynamic Cross-Layer Injection for Deep Vision-Language Fusion

2026-01-15 · Cheng Chen, Yuyu Guo, Pengpeng Zeng, Jingkuan Song 외 arxiv

Vision-Language Models (VLMs) create a severe visual feature bottleneck by using a crude, asymmetric connection that links only the output of the vision encoder to the input of the large language model (LLM). This static…