paper-with-me

Papers

Multi-Explainable TemporalNet: An Interpretable Multimodal Approach using Temporal Convolutional Network for User-level Depression Detection

2024-04-22 · Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition 2024 4 · Anas Zafar, Danyal Aftab, Rizwan Qureshi, Yaofeng Wang, Hong Yan

Multimodal depression detection through internet-based data such as social media platforms has been an important problem in the research community aiming to predict human mental states for ensuring wellbeing of the society. Recently attention-based networks have gained significant popularity for depression detection. However existing multimodal methods primarily rely on images and text assuming no correlation between temporal aspects such as relative time of different posts or tweets which is a crucial factor in deriving depression related behavior patterns. Moreover they lack model interpretability resulting in limited understanding of how different features are contributing to the model's final prediction. In this paper we propose Multi-Explainable TemporalNet (METN) a Temporal Convolution Network (TCN) based multi-modal transformer network with relative timestamp embeddings. We leverage pretrained foundation models for text and image embeddings and attention maps for model interpretability. We perform extensive experiments and ablation studies to validate the performance of METN for user-level depression detection task. Our model shows state-of-the-art results on various benchmarks such as 0.945 F1 score on multimodal Twitter dataset and 0.913 F1 score on multimodal Reddit dataset. We further demonstrate that our model enhances the accuracy of identifying depression in individuals who publicly post messages on social media platforms with enhanced interpretable compatibility.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Depression Detection

Methods 이 논문이 사용한 방법론

Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…

Similar Papers 제목 키워드 기반

TemporalNet: Real-time 2D-3D Video Object Detection

2022-08-29 · Conference on Robots and Vision 2022 8 · Meihong Chen,Jochen Lang

Designing a video detection network based on state-of-the-art single-image object detectors may seem like an obvious choice. However, video object detection has extra challenges due to the lower quality of individual fra…

GPUObjectobject-detectionObject Detection+1

Multi-person Articulated Tracking with Spatial and Temporal Embeddings

2019-03-21 · CVPR 2019 6 · Sheng Jin, Wentao Liu, Wanli Ouyang, Chen Qian

We propose a unified framework for multi-person pose estimation and tracking. Our framework consists of two main components,~\ie~SpatialNet and TemporalNet. The SpatialNet accomplishes body part detection and part-level …

Multi-Object TrackingMulti-Person Pose EstimationMulti-Person Pose Estimation and TrackingObject Tracking+2

GALAR-TemporalNet v2: Anatomy-Guided Dual-Branch Temporal Classification with Bidirectional Mamba and Dual-Graph GCN for Video Capsule Endoscopy -- after competition results

2026-05-21 · Jiye Won, Seangmin Lee, Soon Ki Jung arxiv

Video Capsule Endoscopy (VCE) poses a challenging multi-label temporal classification problem, requiring simultaneous localization of 8 anatomical regions and detection of 9 pathological findings across tens of thousands…

Holmes-VAD: Towards Unbiased and Explainable Video Anomaly Detection via Multi-modal LLM

2024-06-18 · Huaxin Zhang, Xiaohao Xu, Xiang Wang, Jialong Zuo 외

Towards open-ended Video Anomaly Detection (VAD), existing methods often exhibit biased detection when faced with challenging or unseen events and lack interpretability. To address these drawbacks, we propose Holmes-VAD,…

Anomaly DetectionAnomaly LocalizationLanguage ModelingLanguage Modelling+3

Scalable and Explainable Learner-Video Interaction Prediction using Multimodal Large Language Models

2026-04-06 · Dominik Glandorf, Fares Fawzi, Tanja Käser arxiv

Learners' use of video controls in educational videos provides implicit signals of cognitive processing and instructional design quality, yet the lack of scalable and explainable predictive models limits instructors' abi…