LaGeM: A Large Geometry Model for 3D Representation Learning and Diffusion
This paper introduces a novel hierarchical autoencoder that maps 3D models into a highly compressed latent space. The hierarchical autoencoder is specifically designed to tackle the challenges arising from large-scale datasets and generative modeling using diffusion. Different from previous approaches that only work on a regular image or volume grid, our hierarchical autoencoder operates on unordered sets of vectors. Each level of the autoencoder controls different geometric levels of detail. We show that the model can be used to represent a wide range of 3D models while faithfully representing high-resolution geometry details. The training of the new architecture takes 0.70x time and 0.58x memory compared to the baseline. We also explore how the new representation can be used for generative modeling. Specifically, we propose a cascaded diffusion framework where each stage is conditioned on the previous stage. Our design extends existing cascaded designs for image and volume grids to vector sets.
Code (0)
등록된 구현이 없습니다.
Tasks
Representation LearningMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
A Modelagem Computacional do Dom\'\inio dos Esportes na FrameNet Brasil (The Computational Modeling of the Sports Domain in FrameNet Brasil)[In Portuguese]
Descri\cc\~ao e modelagem de constru\cc\~oes interrogativas QU- em Portugu\^es Brasileiro para o desenvolvimento de um chatbot (Description and modeling of interrogative constructs QU- in Brazilian Portuguese for the development of a chatbot)[In Portuguese]
Constru\cc\~oes de Estrutura Argumental no \^ambito do Constructicon da FrameNet Brasil: proposta de uma modelagem lingu\'\istico-computacional (Structural Constructs of Arguments in the Context of the Construction of FrameNet Brasil: a proposal for a computational-linguistic modeling)[In Portuguese]
Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling
Videos inherently represent 2D projections of a dynamic 3D world. However, our analysis suggests that video diffusion models trained solely on raw video data often fail to capture meaningful geometric-aware structure in …
Video GenerationFrom Layers to Networks: Comparing Neural Representations via Diffusion Geometry
Diffusion geometry is a manifold learning framework that uses random walks defined by Markov transition matrices to characterize the geometry of a dataset at multiple scales. We use diffusion geometry for neural represen…