paper-with-me

Papers

LM-MCVT: A Lightweight Multi-modal Multi-view Convolutional-Vision Transformer Approach for 3D Object Recognition

2025-04-27 · Songsong Xiong, Hamidreza Kasaei

In human-centered environments such as restaurants, homes, and warehouses, robots often face challenges in accurately recognizing 3D objects. These challenges stem from the complexity and variability of these environments, including diverse object shapes. In this paper, we propose a novel Lightweight Multi-modal Multi-view Convolutional-Vision Transformer network (LM-MCVT) to enhance 3D object recognition in robotic applications. Our approach leverages the Globally Entropy-based Embeddings Fusion (GEEF) method to integrate multi-views efficiently. The LM-MCVT architecture incorporates pre- and mid-level convolutional encoders and local and global transformers to enhance feature extraction and recognition accuracy. We evaluate our method on the synthetic ModelNet40 dataset and achieve a recognition accuracy of 95.6% using a four-view setup, surpassing existing state-of-the-art methods. To further validate its effectiveness, we conduct 5-fold cross-validation on the real-world OmniObject3D dataset using the same configuration. Results consistently show superior performance, demonstrating the method's robustness in 3D object recognition across synthetic and real-world 3D data.

📄 PDF Abstract BibTeX arXiv:2504.19256

Code (0)

등록된 구현이 없습니다.

Tasks

3D Object RecognitionObjectObject Recognition

Methods 이 논문이 사용한 방법론

Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Multi-Head Attention 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Adam 설명 없음
Attention 설명 없음
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Position-Wise Feed-Forward Layer 설명 없음

Similar Papers 제목 키워드 기반

Edge Assisted Multi-Camera Vehicle Tracking Framework for Real-Time and Scalable Deployment

2025-11-17 · Yuqiang Lin, Sam Lockyer, Shucheng Zhang, Florian Stanek 외 arxiv

Cameras are a core sensing modality in modern intelligent transportation systems (ITS), providing rich visual information on road-user activities. Multi-Camera Vehicle Tracking (MCVT) uses this data to reconstruct vehicl…

Object Detection

RoundaboutHD: High-Resolution Real-World Urban Environment Benchmark for Multi-Camera Vehicle Tracking

2025-07-11 · Yuqiang Lin, Sam Lockyer, Mingxuan Sui, Li Gan 외 arxiv

The multi-camera vehicle tracking (MCVT) framework holds significant potential for smart city applications, including anomaly detection, traffic density estimation, and suspect vehicle tracking. However, current publicly…

Vehicle Re-IdentificationDensity EstimationAnomaly DetectionObject Detection

ELMM: Efficient Lightweight Multimodal Large Language Models for Multimodal Knowledge Graph Completion

2025-10-19 · Wei Huang, Peining Li, Meiyu Liang, Xu Hou 외 arxiv

Multimodal Knowledge Graphs (MKGs) extend traditional knowledge graphs by incorporating visual and textual modalities, enabling richer and more expressive entity representations. However, existing MKGs often suffer from …

Knowledge Graph CompletionKnowledge Graphs

Learning Comprehensive Representations with Richer Self for Text-to-Image Person Re-Identification

2023-10-17 · Shuanglin Yan, Neng Dong, Jun Liu, Liyan Zhang 외

Text-to-image person re-identification (TIReID) retrieves pedestrian images of the same identity based on a query text. However, existing methods for TIReID typically treat it as a one-to-one image-text matching problem,…

Image RetrievalImage-text matchingPerson Re-IdentificationText Matching

SuperLightNet: Lightweight Parameter Aggregation Network for Multimodal Brain Tumor Segmentation

2025-01-01 · CVPR 2025 1 · Feng Yu, Jiacheng Cao, Li Liu, Minghua Jiang

Multimodal 3D segmentation involves a significant number of 3D convolution operations, which requires substantial computational resources and high-performance computing devices in MRI multimodal brain tumor segmentat…

Brain Tumor SegmentationBraTS2021Computational EfficiencyDecoder+3