paper-with-me

Papers

Local Temporal Feature Enhanced Transformer with ROI-rank Based Masking for Diagnosis of ADHD

2025-04-12 · Byunggun Kim, Younghun Kwon

In modern society, Attention-Deficit/Hyperactivity Disorder (ADHD) is one of the common mental diseases discovered not only in children but also in adults. In this context, we propose a ADHD diagnosis transformer model that can effectively simultaneously find important brain spatiotemporal biomarkers from resting-state functional magnetic resonance (rs-fMRI). This model not only learns spatiotemporal individual features but also learns the correlation with full attention structures specialized in ADHD diagnosis. In particular, it focuses on learning local blood oxygenation level dependent (BOLD) signals and distinguishing important regions of interest (ROI) in the brain. Specifically, the three proposed methods for ADHD diagnosis transformer are as follows. First, we design a CNN-based embedding block to obtain more expressive embedding features in brain region attention. It is reconstructed based on the previously CNN-based ADHD diagnosis models for the transformer. Next, for individual spatiotemporal feature attention, we change the attention method to local temporal attention and ROI-rank based masking. For the temporal features of fMRI, the local temporal attention enables to learn local BOLD signal features with only simple window masking. For the spatial feature of fMRI, ROI-rank based masking can distinguish ROIs with high correlation in ROI relationships based on attention scores, thereby providing a more specific biomarker for ADHD diagnosis. The experiment was conducted with various types of transformer models. To evaluate these models, we collected the data from 939 individuals from all sites provided by the ADHD-200 competition. Through this, the spatiotemporal enhanced transformer for ADHD diagnosis outperforms the performance of other different types of transformer variants. (77.78ACC 76.60SPE 79.22SEN 79.30AUC)

📄 PDF Abstract BibTeX arXiv:2504.11474

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

KeyRe-ID: Keypoint-Guided Person Re-Identification using Part-Aware Representation in Videos

2025-07-10 · Jinseong Kim, Jeonghoon Song, Gyeongseon Baek, Byeongjoon Noh

We propose \textbf{KeyRe-ID}, a keypoint-guided video-based person re-identification framework consisting of global and local branches that leverage human keypoints for enhanced spatiotemporal representation learning. Th…

Person Re-IdentificationRepresentation LearningVideo-Based Person Re-Identification

Exploring Stronger Feature for Temporal Action Localization

2021-06-24 · Zhiwu Qing, Xiang Wang, Ziyuan Huang, Yutong Feng 외

Temporal action localization aims to localize starting and ending time with action category. Limited by GPU memory, mainstream methods pre-extract features for each video. Therefore, feature quality determines the upper …

Action LocalizationGPUTemporal Action Localization

Temporal Action Localization with Enhanced Instant Discriminability

2023-09-11 · Dingfeng Shi, Qiong Cao, Yujie Zhong, Shan An 외

Temporal action detection (TAD) aims to detect all action boundaries and their corresponding categories in an untrimmed video. The unclear boundaries of actions in videos often result in imprecise predictions of action b…

Action DetectionAction LocalizationTemporal Action Localization

Multimodal Locally Enhanced Transformer for Continuous Sign Language Recognition

2023-08-22 · Conference of the International Speech Communication Association (INTERSPEECH) 2023 8 · Katerina Papadimitriou, Gerasimos Potamianos

In this paper, we propose a novel Transformer-based approach for continuous sign language recognition (CSLR) from videos, aiming to address the shortcomings of traditional Transformers in learning local semantic context …

Knowledge DistillationPositionSign Language Recognition

CCDSReFormer: Traffic Flow Prediction with a Criss-Crossed Dual-Stream Enhanced Rectified Transformer Model

2024-03-26 · Zhiqi Shao, Michael G. H. Bell, Ze Wang, D. Glenn Geers 외

Accurate, and effective traffic forecasting is vital for smart traffic systems, crucial in urban traffic planning and management. Current Spatio-Temporal Transformer models, despite their prediction capabilities, struggl…

Computational EfficiencyManagement