Multi-Explainable TemporalNet: An Interpretable Multimodal Approach using Temporal Convolutional Network for User-level Depression Detection
Multimodal depression detection through internet-based data such as social media platforms has been an important problem in the research community aiming to predict human mental states for ensuring wellbeing of the society. Recently attention-based networks have gained significant popularity for depression detection. However existing multimodal methods primarily rely on images and text assuming no correlation between temporal aspects such as relative time of different posts or tweets which is a crucial factor in deriving depression related behavior patterns. Moreover they lack model interpretability resulting in limited understanding of how different features are contributing to the model's final prediction. In this paper we propose Multi-Explainable TemporalNet (METN) a Temporal Convolution Network (TCN) based multi-modal transformer network with relative timestamp embeddings. We leverage pretrained foundation models for text and image embeddings and attention maps for model interpretability. We perform extensive experiments and ablation studies to validate the performance of METN for user-level depression detection task. Our model shows state-of-the-art results on various benchmarks such as 0.945 F1 score on multimodal Twitter dataset and 0.913 F1 score on multimodal Reddit dataset. We further demonstrate that our model enhances the accuracy of identifying depression in individuals who publicly post messages on social media platforms with enhanced interpretable compatibility.
Code (0)
등록된 구현이 없습니다.
Tasks
Depression DetectionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
TemporalNet: Real-time 2D-3D Video Object Detection
Designing a video detection network based on state-of-the-art single-image object detectors may seem like an obvious choice. However, video object detection has extra challenges due to the lower quality of individual fra…
GPUObjectobject-detectionObject Detection+1Multi-person Articulated Tracking with Spatial and Temporal Embeddings
We propose a unified framework for multi-person pose estimation and tracking. Our framework consists of two main components,~\ie~SpatialNet and TemporalNet. The SpatialNet accomplishes body part detection and part-level …
Multi-Object TrackingMulti-Person Pose EstimationMulti-Person Pose Estimation and TrackingObject Tracking+2GALAR-TemporalNet v2: Anatomy-Guided Dual-Branch Temporal Classification with Bidirectional Mamba and Dual-Graph GCN for Video Capsule Endoscopy -- after competition results
Video Capsule Endoscopy (VCE) poses a challenging multi-label temporal classification problem, requiring simultaneous localization of 8 anatomical regions and detection of 9 pathological findings across tens of thousands…
Holmes-VAD: Towards Unbiased and Explainable Video Anomaly Detection via Multi-modal LLM
Towards open-ended Video Anomaly Detection (VAD), existing methods often exhibit biased detection when faced with challenging or unseen events and lack interpretability. To address these drawbacks, we propose Holmes-VAD,…
Anomaly DetectionAnomaly LocalizationLanguage ModelingLanguage Modelling+3Scalable and Explainable Learner-Video Interaction Prediction using Multimodal Large Language Models
Learners' use of video controls in educational videos provides implicit signals of cognitive processing and instructional design quality, yet the lack of scalable and explainable predictive models limits instructors' abi…