paper-with-me

홈 › Papers

UniMesh: Unifying 3D Mesh Understanding and Generation

2026-04-19 · Peng Huang, Yifeng Chen, Zeyu Zhang, Hao Tang arxiv

Recent advances in 3D vision have led to specialized models for either 3D understanding (e.g., shape classification, segmentation, reconstruction) or 3D generation (e.g., synthesis, completion, and editing). However, these tasks are often tackled in isolation, resulting in fragmented architectures and representations that hinder knowledge transfer and holistic scene modeling. To address these challenges, we propose UniMesh, a unified framework that jointly learns 3D generation and understanding within a single architecture. First, we introduce a novel Mesh Head that acts as a cross model interface, bridging diffusion based image generation with implicit shape decoders. Second, we develop Chain of Mesh (CoM), a geometric instantiation of iterative reasoning that enables user driven semantic mesh editing through a closed loop latent, prompting, and re generation cycle. Third, we incorporate a self reflection mechanism based on an Actor Evaluator Self reflection triad to diagnose and correct failures in high level tasks like 3D captioning. Experimental results demonstrate that UniMesh not only achieves competitive performance on standard benchmarks but also unlocks novel capabilities in iterative editing and mutual enhancement between generation and understanding. Code: https://github.com/AIGeeksGroup/UniMesh. Website: https://aigeeksgroup.github.io/UniMesh.

📄 PDF Abstract BibTeX arXiv:2604.17472

Code (0)

등록된 구현이 없습니다.

Tasks

Image Generation3D Generation

Similar Papers 제목 키워드 기반

LLaMA-Mesh: Unifying 3D Mesh Generation with Language Models

2024-11-14 · Zhengyi Wang, Jonathan Lorraine, Yikai Wang, Hang Su 외

This work explores expanding the capabilities of large language models (LLMs) pretrained on text to generate 3D meshes within a unified model. This offers key advantages of (1) leveraging spatial knowledge already embedd…

3D GenerationText Generation

PolyFlow: Continuous Topology Embedding Flow Matching for Artist-style Mesh Generation

2026-06-25 · Chunshi Wang, Haohan Weng, Junliang Ye, Biwen Lei 외 hf

Autoregressive Transformers dominate high-quality mesh generation by producing artist-worthy topologies, yet their inherent sequential decoding induces substantial computational overhead, falling orders of magnitude slow…

Towards Unifying Understanding and Generation in the Era of Vision Foundation Models: A Survey from the Autoregression Perspective

2024-10-29 · Shenghao Xie, Wenqiang Zu, Mingyang Zhao, Duo Su 외

Autoregression in large language models (LLMs) has shown impressive scalability by unifying all language tasks into the next token prediction paradigm. Recently, there is a growing interest in extending this success to v…

Survey

MeshPose: Unifying DensePose and 3D Body Mesh reconstruction

2024-06-14 · CVPR 2024 1 · Eric-Tuan Lê, Antonis Kakolyris, Petros Koutras, Himmy Tam 외

DensePose provides a pixel-accurate association of images with 3D mesh coordinates, but does not provide a 3D mesh, while Human Mesh Reconstruction (HMR) systems have high 2D reprojection error, as measured by DensePose …

3D Human Pose EstimationPose Estimation

A Simple Baseline for Unifying Understanding, Generation, and Editing via Vanilla Next-token Prediction

2026-03-05 · Jie Zhu, Hanghang Ma, Jia Wang, Yayong Guan 외 arxiv

In this work, we introduce Wallaroo, a simple autoregressive baseline that leverages next-token prediction to unify multi-modal understanding, image generation, and editing at the same time. Moreover, Wallaroo supports m…

Image Generation