paper-with-me

Papers

E-M3RF: An Equivariant Multimodal 3D Re-assembly Framework

2025-11-26 · Adeela Islam, Stefano Fiorini, Manuel Lecha, Theodore Tsesmelis, Stuart James, Pietro Morerio, Alessio Del Bue arxiv

3D reassembly is a fundamental geometric problem, and in recent years it has increasingly been challenged by deep learning methods rather than classical optimization. While learning approaches have shown promising results, most still rely primarily on geometric features to assemble a whole from its parts. As a result, methods struggle when geometry alone is insufficient or ambiguous, for example, for small, eroded, or symmetric fragments. Additionally, solutions do not impose physical constraints that explicitly prevent overlapping assemblies. To address these limitations, we introduce E-M3RF, an equivariant multimodal 3D reassembly framework that takes as input the point clouds, containing both point positions and colors of fractured fragments, and predicts the transformations required to reassemble them using SE(3) flow matching. Each fragment is represented by both geometric and color features: i) 3D point positions are encoded as rotationconsistent geometric features using a rotation-equivariant encoder, ii) the colors at each 3D point are encoded with a transformer. The two feature sets are then combined to form a multimodal representation. We experimented on four datasets: two synthetic datasets, Breaking Bad and Fantastic Breaks, and two real-world cultural heritage datasets, RePAIR and Presious, demonstrating that E-M3RF on the RePAIR dataset reduces rotation error by 23.1% and translation error by 13.2%, while Chamfer Distance decreases by 18.4% compared to competing methods.

📄 PDF Abstract BibTeX arXiv:2511.21422

Code (0)

등록된 구현이 없습니다.

Tasks

Point Clouds

Similar Papers 제목 키워드 기반

LEGO Co-builder: Exploring Fine-Grained Vision-Language Modeling for Multimodal LEGO Assembly Assistants

2025-07-07 · Haochen Huang, Jiahuan Pei, Mohammad Aliannejadi, Xin Sun 외 arxiv

Vision-language models (VLMs) are facing the challenges of understanding and following multimodal assembly instructions, particularly when fine-grained spatial reasoning and precise object state detection are required. I…

Spatial ReasoningObject Detection

ProMQA-Assembly: Multimodal Procedural QA Dataset on Assembly

2025-09-03 · Kimihiro Hasegawa, Wiradee Imrattanatrai, Masaki Asada, Susan Holm 외 arxiv

Assistants on assembly tasks show great potential to benefit humans ranging from helping with everyday tasks to interacting in industrial settings. However, evaluation resources in assembly activities are underexplored. …

AssemblyBench: Physics-Aware Assembly of Complex Industrial Objects

2026-05-13 · Danrui Li, Jiahao Zhang, Bernhard Egger, Moitreya Chatterjee 외 arxiv

Assembling objects from parts requires understanding multimodal instructions, linking them to 3D components, and predicting physically plausible 6-DoF motions for each assembly step. Existing datasets focus on simplified…

Pose Estimation

Combinative Matching for Geometric Shape Assembly

2025-08-13 · Nahyuk Lee, Juhong Min, Junhong Lee, Chunghyun Park 외 arxiv

This paper introduces a new shape-matching methodology, combinative matching, to combine interlocking parts for geometric shape assembly. Previous methods for geometric assembly typically rely on aligning parts by findin…

Two by Two: Learning Multi-Task Pairwise Objects Assembly for Generalizable Robot Manipulation

2025-04-09 · CVPR 2025 1 · Yu Qi, Yuanchen Ju, Tianming Wei, Chi Chu 외

3D assembly tasks, such as furniture assembly and component fitting, play a crucial role in daily life and represent essential capabilities for future home robots. Existing benchmarks and datasets predominantly focus on …

3D AssemblyPose EstimationRobot Manipulation