Papers Multi-modal Classification
“Multi-modal Classification” 태그가 달린 논문 37편 · 필터 해제
Multi-Modal Building Inspection via Perceiver IO Fusion of Satellite and Street-Level Imagery
We present a multi-modal classification framework that fuses satellite and street-level imagery through a Perceiver IO architecture operating on spatial patch tokens from a shared DINOv2 backbone. The design naturally ha…
Multi-modal ClassificationA Hybrid CNN and ML Framework for Multi-modal Classification of Movement Disorders Using MRI and Brain Structural Features
Atypical Parkinsonian Disorders (APD), also known as Parkinson-plus syndrome, are a group of neurodegenerative diseases that include progressive supranuclear palsy (PSP) and multiple system atrophy (MSA). In the early st…
Multi-modal ClassificationToken Entropy Regularization for Multi-modal Antenna Affiliation Identification
Accurate antenna affiliation identification is crucial for optimizing and maintaining communication networks. Current practice, however, relies on the cumbersome and error-prone process of manual tower inspections. We pr…
Multi-modal ClassificationD-CAT: Decoupled Cross-Attention Transfer between Sensor Modalities for Unimodal Inference
Cross-modal transfer learning is used to improve multi-modal classification models (e.g., for human activity recognition in human-robot collaboration). However, existing methods require paired sensor data at both trainin…
Human Activity RecognitionMulti-modal ClassificationTransfer LearningSurformer v2: A Multimodal Classifier for Surface Understanding from Touch and Vision
Multimodal surface material classification plays a critical role in advancing tactile perception for robotic manipulation and interaction. In this paper, we present Surformer v2, an enhanced multi-modal classification ar…
Multi-modal ClassificationBi-cephalic self-attended model to classify Parkinson's disease patients with freezing of gait
Parkinson's Disease (PD) often results in motor and cognitive impairments, including gait dysfunction, particularly in patients with freezing of gait (FOG). Current detection methods are either subjective or reliant on s…
Multi-modal ClassificationLightweight Joint Audio-Visual Deepfake Detection via Single-Stream Multi-Modal Learning Framework
Deepfakes are AI-synthesized multimedia data that may be abused for spreading misinformation. Deepfake generation involves both visual and audio manipulation. To detect audio-visual deepfakes, previous studies commonly e…
audio-visual learningDeepFake DetectionFace SwappingMisinformation+1A Survey on Training-free Open-Vocabulary Semantic Segmentation
Semantic segmentation is one of the most fundamental tasks in image understanding with a long history of research, and subsequently a myriad of different approaches. Traditional methods strive to train models up from scr…
Multi-modal ClassificationOpen Vocabulary Semantic SegmentationOpen-Vocabulary Semantic SegmentationSemantic Segmentation+1A Comparative Study of Human Activity Recognition: Motion, Tactile, and multi-modal Approaches
Human activity recognition (HAR) is essential for effective Human-Robot Collaboration (HRC), enabling robots to interpret and respond to human actions. This study evaluates the ability of a vision-based tactile sensor to…
Activity RecognitionClassificationHuman Activity RecognitionMulti-modal ClassificationMulti-modal classification of forest biodiversity potential from 2D orthophotos and 3D airborne laser scanning point clouds
Accurate assessment of forest biodiversity is crucial for ecosystem management and conservation. While traditional field surveys provide high-quality assessments, they are labor-intensive and spatially limited. This stud…
Multi-modal ClassificationMultimodal Learning with Uncertainty Quantification based on Discounted Belief Fusion
Multimodal AI models are increasingly used in fields like healthcare, finance, and autonomous driving, where information is drawn from multiple sources or modalities such as images, texts, audios, videos. However, effect…
Decision MakingMulti-modal ClassificationUncertainty QuantificationHateful Meme Detection through Context-Sensitive Prompting and Fine-Grained Labeling
The prevalence of multi-modal content on social media complicates automated moderation strategies. This calls for an enhancement in multi-modal classification and a deeper understanding of understated meanings in images …
Model OptimizationMulti-modal ClassificationTurbo your multi-modal classification with contrastive learning
Contrastive learning has become one of the most impressive approaches for multi-modal representation learning. However, previous multi-modal works mainly focused on cross-modal understanding, ignoring in-modal contrastiv…
ClassificationContrastive LearningEmotion RecognitionMulti-modal Classification+4FungiTastic: A multi-modal dataset and benchmark for image categorization
We introduce a new, challenging benchmark and a dataset, FungiTastic, based on fungal records continuously collected over a twenty-year span. The dataset is labeled and curated by experts and consists of about 350k multi…
ClassificationFew-Shot LearningImage CategorizationMulti-modal Classification+1Language Augmentation in CLIP for Improved Anatomy Detection on Multi-modal Medical Images
Vision-language models have emerged as a powerful tool for previously challenging multi-modal classification problem in the medical domain. This development has led to the exploration of automated image description gener…
AnatomyImage DescriptionMulti-modal ClassificationJoint-Individual Fusion Structure with Fusion Attention Module for Multi-Modal Skin Cancer Classification
Most convolutional neural network (CNN) based methods for skin cancer classification obtain their results using only dermatological images. Although good classification results have been shown, more accurate results can …
Cancer ClassificationClassificationDecision MakingMulti-modal Classification+1PromptStyler: Prompt-driven Style Generation for Source-free Domain Generalization
In a joint vision-language space, a text feature (e.g., from "a photo of a dog") could effectively represent its relevant image features (e.g., from dog photos). Also, a recent study has demonstrated the cross-modal tran…
Domain GeneralizationImage ClassificationMulti-modal ClassificationMultimodal Deep Learning+4FAME-ViL: Multi-Tasking Vision-Language Model for Heterogeneous Fashion Tasks
In the fashion domain, there exists a variety of vision-and-language (V+L) tasks, including cross-modal retrieval, text-guided image retrieval, multi-modal classification, and image captioning. They differ drastically in…
Cross-Modal RetrievalImage CaptioningImage RetrievalLanguage Modeling+3Contrastive Audio-Visual Masked Autoencoder
In this paper, we first extend the recent Masked Auto-Encoder (MAE) model from a single modality to audio-visual multi-modalities. Subsequently, we propose the Contrastive Audio-Visual Masked Auto-Encoder (CAV-MAE) by co…
Audio ClassificationAudio TaggingContrastive LearningMulti-modal Classification+4AVT: Audio-Video Transformer for Multimodal Action Recognition
Action recognition is an essential field for video understanding. To learn from heterogeneous data sources effectively, in this work, we propose a novel multimodal action recognition approach termed Audio-Video Transform…
Action RecognitionAudio ClassificationContrastive LearningMulti-modal Classification+1