paper-with-me

Papers Multi-modal Classification

“Multi-modal Classification” 태그가 달린 논문 37편 · 필터 해제

Multi-Modal Building Inspection via Perceiver IO Fusion of Satellite and Street-Level Imagery

2026-05-25 · Niels Sombekke, Rob G. J. Wijnhoven, Martin R. Oswald arxiv

We present a multi-modal classification framework that fuses satellite and street-level imagery through a Perceiver IO architecture operating on spatial patch tokens from a shared DINOv2 backbone. The design naturally ha…

Multi-modal Classification

A Hybrid CNN and ML Framework for Multi-modal Classification of Movement Disorders Using MRI and Brain Structural Features

2026-02-05 · Mengyu Li, Ingibjörg Kristjánsdóttir, Thilo van Eimeren, Kathrin Giehl 외 arxiv

Atypical Parkinsonian Disorders (APD), also known as Parkinson-plus syndrome, are a group of neurodegenerative diseases that include progressive supranuclear palsy (PSP) and multiple system atrophy (MSA). In the early st…

Multi-modal Classification

Token Entropy Regularization for Multi-modal Antenna Affiliation Identification

2026-01-29 · Dong Chen, Ruoyu Li, Xinyan Zhang, Jialei Xu 외 arxiv

Accurate antenna affiliation identification is crucial for optimizing and maintaining communication networks. Current practice, however, relies on the cumbersome and error-prone process of manual tower inspections. We pr…

Multi-modal Classification

D-CAT: Decoupled Cross-Attention Transfer between Sensor Modalities for Unimodal Inference

2025-09-11 · Leen Daher, Zhaobo Wang, Malcolm Mielle arxiv

Cross-modal transfer learning is used to improve multi-modal classification models (e.g., for human activity recognition in human-robot collaboration). However, existing methods require paired sensor data at both trainin…

Human Activity RecognitionMulti-modal ClassificationTransfer Learning

Surformer v2: A Multimodal Classifier for Surface Understanding from Touch and Vision

2025-09-04 · Manish Kansana, Sindhuja Penchala, Shahram Rahimi, Noorbakhsh Amiri Golilarz arxiv

Multimodal surface material classification plays a critical role in advancing tactile perception for robotic manipulation and interaction. In this paper, we present Surformer v2, an enhanced multi-modal classification ar…

Multi-modal Classification

Bi-cephalic self-attended model to classify Parkinson's disease patients with freezing of gait

2025-07-28 · Shomoita Jahid Mitin, Rodrigue Rizk, Maximilian Scherer, Thomas Koeglsperger 외 arxiv

Parkinson's Disease (PD) often results in motor and cognitive impairments, including gait dysfunction, particularly in patients with freezing of gait (FOG). Current detection methods are either subjective or reliant on s…

Multi-modal Classification

Lightweight Joint Audio-Visual Deepfake Detection via Single-Stream Multi-Modal Learning Framework

2025-06-09 · Kuiyuan Zhang, Wenjie Pei, Rushi Lan, Yifang Guo 외

Deepfakes are AI-synthesized multimedia data that may be abused for spreading misinformation. Deepfake generation involves both visual and audio manipulation. To detect audio-visual deepfakes, previous studies commonly e…

audio-visual learningDeepFake DetectionFace SwappingMisinformation+1

A Survey on Training-free Open-Vocabulary Semantic Segmentation

2025-05-28 · Naomi Kombol, Ivan Martinović, Siniša Šegvić

Semantic segmentation is one of the most fundamental tasks in image understanding with a long history of research, and subsequently a myriad of different approaches. Traditional methods strive to train models up from scr…

Multi-modal ClassificationOpen Vocabulary Semantic SegmentationOpen-Vocabulary Semantic SegmentationSemantic Segmentation+1

A Comparative Study of Human Activity Recognition: Motion, Tactile, and multi-modal Approaches

2025-05-13 · Valerio Belcamino, Nhat Minh Dinh Le, Quan Khanh Luu, Alessandro Carfì 외

Human activity recognition (HAR) is essential for effective Human-Robot Collaboration (HRC), enabling robots to interpret and respond to human actions. This study evaluates the ability of a vision-based tactile sensor to…

Activity RecognitionClassificationHuman Activity RecognitionMulti-modal Classification

Multi-modal classification of forest biodiversity potential from 2D orthophotos and 3D airborne laser scanning point clouds

2025-01-03 · Simon B. Jensen, Stefan Oehmcke, Andreas Møgelmose, Meysam Madadi 외

Accurate assessment of forest biodiversity is crucial for ecosystem management and conservation. While traditional field surveys provide high-quality assessments, they are labor-intensive and spatially limited. This stud…

Multi-modal Classification

Multimodal Learning with Uncertainty Quantification based on Discounted Belief Fusion

2024-12-23 · Grigor Bezirganyan, Sana Sellami, Laure Berti-ÉQuille, Sébastien Fournier

Multimodal AI models are increasingly used in fields like healthcare, finance, and autonomous driving, where information is drawn from multiple sources or modalities such as images, texts, audios, videos. However, effect…

Decision MakingMulti-modal ClassificationUncertainty Quantification

Hateful Meme Detection through Context-Sensitive Prompting and Fine-Grained Labeling

2024-11-13 · Rongxin Ouyang, Kokil Jaidka, Subhayan Mukerjee, Guangyu Cui

The prevalence of multi-modal content on social media complicates automated moderation strategies. This calls for an enhancement in multi-modal classification and a deeper understanding of understated meanings in images …

Model OptimizationMulti-modal Classification

Turbo your multi-modal classification with contrastive learning

2024-09-14 · ZhiYu Zhang, Da Liu, Shengqiang Liu, Anna Wang 외

Contrastive learning has become one of the most impressive approaches for multi-modal representation learning. However, previous multi-modal works mainly focused on cross-modal understanding, ignoring in-modal contrastiv…

ClassificationContrastive LearningEmotion RecognitionMulti-modal Classification+4

FungiTastic: A multi-modal dataset and benchmark for image categorization

2024-08-24 · Lukas Picek, Klara Janouskova, Milan Sulc, Jiri Matas

We introduce a new, challenging benchmark and a dataset, FungiTastic, based on fungal records continuously collected over a twenty-year span. The dataset is labeled and curated by experts and consists of about 350k multi…

ClassificationFew-Shot LearningImage CategorizationMulti-modal Classification+1

Language Augmentation in CLIP for Improved Anatomy Detection on Multi-modal Medical Images

2024-05-31 · Mansi Kakkar, Dattesh Shanbhag, Chandan Aladahalli, Gurunath Reddy M

Vision-language models have emerged as a powerful tool for previously challenging multi-modal classification problem in the medical domain. This development has led to the exploration of automated image description gener…

AnatomyImage DescriptionMulti-modal Classification

Joint-Individual Fusion Structure with Fusion Attention Module for Multi-Modal Skin Cancer Classification

2023-12-07 · Peng Tang, Xintong Yan, Yang Nan, Xiaobin Hu 외

Most convolutional neural network (CNN) based methods for skin cancer classification obtain their results using only dermatological images. Although good classification results have been shown, more accurate results can …

Cancer ClassificationClassificationDecision MakingMulti-modal Classification+1

PromptStyler: Prompt-driven Style Generation for Source-free Domain Generalization

2023-07-27 · ICCV 2023 1 · Junhyeong Cho, Gilhyun Nam, Sungyeon Kim, Hunmin Yang 외

In a joint vision-language space, a text feature (e.g., from "a photo of a dog") could effectively represent its relevant image features (e.g., from dog photos). Also, a recent study has demonstrated the cross-modal tran…

Domain GeneralizationImage ClassificationMulti-modal ClassificationMultimodal Deep Learning+4

FAME-ViL: Multi-Tasking Vision-Language Model for Heterogeneous Fashion Tasks

2023-03-04 · CVPR 2023 1 · Xiao Han, Xiatian Zhu, Licheng Yu, Li Zhang 외

In the fashion domain, there exists a variety of vision-and-language (V+L) tasks, including cross-modal retrieval, text-guided image retrieval, multi-modal classification, and image captioning. They differ drastically in…

Cross-Modal RetrievalImage CaptioningImage RetrievalLanguage Modeling+3

Contrastive Audio-Visual Masked Autoencoder

2022-10-02 · Yuan Gong, Andrew Rouditchenko, Alexander H. Liu, David Harwath 외

In this paper, we first extend the recent Masked Auto-Encoder (MAE) model from a single modality to audio-visual multi-modalities. Subsequently, we propose the Contrastive Audio-Visual Masked Auto-Encoder (CAV-MAE) by co…

Audio ClassificationAudio TaggingContrastive LearningMulti-modal Classification+4

AVT: Audio-Video Transformer for Multimodal Action Recognition

2022-09-22 · Submitted to ICLR 2022 9 · Wentao Zhu, Jingru Yi, Kevin Hsu, Xiaohang Sun 외

Action recognition is an essential field for video understanding. To learn from heterogeneous data sources effectively, in this work, we propose a novel multimodal action recognition approach termed Audio-Video Transform…

Action RecognitionAudio ClassificationContrastive LearningMulti-modal Classification+1
1–20 / 37 다음 →