paper-with-me

Papers

PatentNet: A Large-Scale Incomplete Multiview, Multimodal, Multilabel Industrial Goods Image Database

2021-06-23 · FangYuan Lei, Da Huang, Jianjian Jiang, Ruijun Ma, Senhong Wang, Jiangzhong Cao, Yusen Lin, Qingyun Dai

In deep learning area, large-scale image datasets bring a breakthrough in the success of object recognition and retrieval. Nowadays, as the embodiment of innovation, the diversity of the industrial goods is significantly larger, in which the incomplete multiview, multimodal and multilabel are different from the traditional dataset. In this paper, we introduce an industrial goods dataset, namely PatentNet, with numerous highly diverse, accurate and detailed annotations of industrial goods images, and corresponding texts. In PatentNet, the images and texts are sourced from design patent. Within over 6M images and corresponding texts of industrial goods labeled manually checked by professionals, PatentNet is the first ongoing industrial goods image database whose varieties are wider than industrial goods datasets used previously for benchmarking. PatentNet organizes millions of images into 32 classes and 219 subclasses based on the Locarno Classification Agreement. Through extensive experiments on image classification, image retrieval and incomplete multiview clustering, we demonstrate that our PatentNet is much more diverse, complex, and challenging, enjoying higher potentials than existing industrial image datasets. Furthermore, the characteristics of incomplete multiview, multimodal and multilabel in PatentNet are able to offer unparalleled opportunities in the artificial intelligence community and beyond.

📄 PDF Abstract BibTeX arXiv:2106.12139

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingClusteringDiversityimage-classificationImage ClassificationImage RetrievalMultiview ClusteringObject RecognitionRetrieval

Similar Papers 제목 키워드 기반

Interpretable Multimodal Cancer Prototyping with Whole Slide Images and Incompletely Paired Genomics

2025-11-26 · Yupei Zhang, Yating Huang, Wanming Hu, Lequan Yu 외 arxiv

Multimodal approaches that integrate histology and genomics hold strong potential for precision oncology. However, phenotypic and genotypic heterogeneity limits the quality of intra-modal representations and hinders effe…

Consistency-Aware Padding for Incomplete Multi-Modal Alignment Clustering Based on Self-Repellent Greedy Anchor Search

2025-07-05 · Shubin Ma, Liang Zhao, Mingdong Lu, Yifan Guo 외 arxiv

Multimodal representation is faithful and highly effective in describing real-world data samples' characteristics by describing their complementary information. However, the collected data often exhibits incomplete and m…

Contrastive Learning

Learning multiview embeddings for assessing dementia

2018-10-01 · EMNLP 2018 10 · Chlo{\'e} Pou-Prom, Frank Rudzicz

As the incidence of Alzheimer{'}s Disease (AD) increases, early detection becomes crucial. Unfortunately, datasets for AD assessment are often sparse and incomplete. In this work, we leverage the multiview nature of a sm…

General Classificationregression

Emotion-LLaMAv2 and MMEVerse: A New Framework and Benchmark for Multimodal Emotion Understanding

2026-01-23 · Xiaojiang Peng, Jingyi Chen, Zebang Cheng, Bao Peng 외 arxiv

Understanding human emotions from multimodal signals poses a significant challenge in affective computing and human-robot interaction. While multimodal large language models (MLLMs) have excelled in general vision-langua…

Emotion RecognitionFace Detection

SiM3D: Single-instance Multiview Multimodal and Multisetup 3D Anomaly Detection Benchmark

2025-06-26 · Alex Costanzino, Pierluigi Zama Ramirez, Luigi Lella, Matteo Ragaglia 외

We propose SiM3D, the first benchmark considering the integration of multiview and multimodal information for comprehensive 3D anomaly detection and segmentation (ADS), where the task is to produce a voxel-based Anomaly …

3D Anomaly Detection3D Anomaly Detection and SegmentationAnomaly Detection