paper-with-me

홈 › Papers

SIM-Net: A Multimodal Fusion Network Using Inferred 3D Object Shape Point Clouds from RGB Images for 2D Classification

2025-06-23 · Youcef Sklab, Hanane Ariouat, Eric Chenin, Edi Prifti, Jean-Daniel Zucker

We introduce the Shape-Image Multimodal Network (SIM-Net), a novel 2D image classification architecture that integrates 3D point cloud representations inferred directly from RGB images. Our key contribution lies in a pixel-to-point transformation that converts 2D object masks into 3D point clouds, enabling the fusion of texture-based and geometric features for enhanced classification performance. SIM-Net is particularly well-suited for the classification of digitized herbarium specimens (a task made challenging by heterogeneous backgrounds), non-plant elements, and occlusions that compromise conventional image-based models. To address these issues, SIM-Net employs a segmentation-based preprocessing step to extract object masks prior to 3D point cloud generation. The architecture comprises a CNN encoder for 2D image features and a PointNet-based encoder for geometric features, which are fused into a unified latent space. Experimental evaluations on herbarium datasets demonstrate that SIM-Net consistently outperforms ResNet101, achieving gains of up to 9.9% in accuracy and 12.3% in F-score. It also surpasses several transformer-based state-of-the-art architectures, highlighting the benefits of incorporating 3D structural reasoning into 2D image classification tasks.

📄 PDF Abstract BibTeX arXiv:2506.18683

Code (0)

등록된 구현이 없습니다.

Tasks

Classificationimage-classificationImage ClassificationPoint Cloud Generation

Similar Papers 제목 키워드 기반

DC3DO: Diffusion Classifier for 3D Objects

2024-08-13 · Nursena Koprucu, Meher Shashwat Nigam, Shicheng Xu, Biruk Abere 외

Inspired by Geoffrey Hinton emphasis on generative modeling, To recognize shapes, first learn to generate them, we explore the use of 3D diffusion models for object classification. Leveraging the density estimates from t…

3D Object ClassificationClassificationMultimodal ReasoningObject+2

LLM-Guided Material Inference for 3D Point Clouds

2025-12-02 · Nafiseh Izadyar, Teseo Schneider arxiv

Most existing 3D shape datasets and models focus solely on geometry, overlooking the material properties that determine how objects appear. We introduce a two-stage large language model (LLM) based method for inferring m…

Point Clouds

Variational Shape Inference for Grasp Diffusion on SE(3)

2025-08-24 · S. Talha Bukhari, Kaivalya Agrawal, Zachary Kingston, Aniket Bera arxiv

Grasp synthesis is a fundamental task in robotic manipulation which usually has multiple feasible solutions. Multimodal grasp synthesis seeks to generate diverse sets of stable grasps conditioned on object geometry, maki…

RealDiff: Real-world 3D Shape Completion using Self-Supervised Diffusion Models

2024-09-16 · Başak Melis Öcal, Maxim Tatarchenko, Sezer Karaoglu, Theo Gevers

Point cloud completion aims to recover the complete 3D shape of an object from partial observations. While approaches relying on synthetic shape priors achieved promising results in this domain, their applicability and g…

ObjectPoint Cloud Completion

MANet: Multimodal Attention Network based Point- View fusion for 3D Shape Recognition

2020-02-28 · Yaxin Zhao, Jichao Jiao, Tangkun Zhang

3D shape recognition has attracted more and more attention as a task of 3D vision research. The proliferation of 3D data encourages various deep learning methods based on 3D data. Now there have been many deep learning m…

3D Shape Recognition