paper-with-me

홈 › Papers

ViSketch-GPT: Collaborative Multi-Scale Feature Extraction for Sketch Recognition and Generation

2025-03-28 · Giulio Federico, Giuseppe Amato, Fabio Carrara, Claudio Gennaro, Marco Di Benedetto

Understanding the nature of human sketches is challenging because of the wide variation in how they are created. Recognizing complex structural patterns improves both the accuracy in recognizing sketches and the fidelity of the generated sketches. In this work, we introduce ViSketch-GPT, a novel algorithm designed to address these challenges through a multi-scale context extraction approach. The model captures intricate details at multiple scales and combines them using an ensemble-like mechanism, where the extracted features work collaboratively to enhance the recognition and generation of key details crucial for classification and generation tasks. The effectiveness of ViSketch-GPT is validated through extensive experiments on the QuickDraw dataset. Our model establishes a new benchmark, significantly outperforming existing methods in both classification and generation tasks, with substantial improvements in accuracy and the fidelity of generated sketches. The proposed algorithm offers a robust framework for understanding complex structures by extracting features that collaborate to recognize intricate details, enhancing the understanding of structures like sketches and making it a versatile tool for various applications in computer vision and machine learning.

📄 PDF Abstract BibTeX arXiv:2503.22374

Code (0)

등록된 구현이 없습니다.

Tasks

Sketch Recognition

Similar Papers 제목 키워드 기반

Building-road Collaborative Extraction from Remotely Sensed Images via Cross-Interaction

2023-07-23 · HaoNan Guo, Xin Su, Chen Wu, Bo Du 외

Buildings are the basic carrier of social production and human life; roads are the links that interconnect social networks. Building and road information has important application value in the frontier fields of regional…

CADA: Multi-scale Collaborative Adversarial Domain Adaptation for Unsupervised Optic Disc and Cup Segmentation

2021-10-05 · Peng Liu, Charlie T. Tran, Bin Kong, Ruogu Fang

The diversity of retinal imaging devices poses a significant challenge: domain shift, which leads to performance degradation when applying the deep learning models trained on one domain to new testing domains. In this pa…

Domain AdaptationUnsupervised Domain Adaptation

Multi-local Collaborative AutoEncoder

2019-06-12 · Jielei Chu, Hongjun Wang, Jing Liu, Zhiguo Gong 외

The excellent performance of representation learning of autoencoders have attracted considerable interest in various applications. However, the structure and multi-local collaborative relationships of unlabeled data are …

ClusteringRepresentation Learning

Deep Flow Collaborative Network for Online Visual Tracking

2019-11-05 · Peidong Liu, Xiyu Yan, Yong Jiang, Shu-Tao Xia

The deep learning-based visual tracking algorithms such as MDNet achieve high performance leveraging to the feature extraction ability of a deep neural network. However, the tracking efficiency of these trackers is not v…

Optical Flow EstimationSchedulingVisual Tracking

MsaMIL-Net: An End-to-End Multi-Scale Aware Multiple Instance Learning Network for Efficient Whole Slide Image Classification

2025-03-11 · Jiangping Wen, Jinyu Wen, Meie Fang

Bag-based Multiple Instance Learning (MIL) approaches have emerged as the mainstream methodology for Whole Slide Image (WSI) classification. However, most existing methods adopt a segmented training strategy, which first…

image-classificationImage ClassificationMultiple Instance Learning