paper-with-me

홈 › Papers

STROKEVISION-BENCH: A Multimodal Video And 2D Pose Benchmark For Tracking Stroke Recovery

2025-09-02 · David Robinson, Animesh Gupta, Rizwan Quershi, Qiushi Fu, Mubarak Shah arxiv

Despite advancements in rehabilitation protocols, clinical assessment of upper extremity (UE) function after stroke largely remains subjective, relying heavily on therapist observation and coarse scoring systems. This subjectivity limits the sensitivity of assessments to detect subtle motor improvements, which are critical for personalized rehabilitation planning. Recent progress in computer vision offers promising avenues for enabling objective, quantitative, and scalable assessment of UE motor function. Among standardized tests, the Box and Block Test (BBT) is widely utilized for measuring gross manual dexterity and tracking stroke recovery, providing a structured setting that lends itself well to computational analysis. However, existing datasets targeting stroke rehabilitation primarily focus on daily living activities and often fail to capture clinically structured assessments such as block transfer tasks. Furthermore, many available datasets include a mixture of healthy and stroke-affected individuals, limiting their specificity and clinical utility. To address these critical gaps, we introduce StrokeVision-Bench, the first-ever dedicated dataset of stroke patients performing clinically structured block transfer tasks. StrokeVision-Bench comprises 1,000 annotated videos categorized into four clinically meaningful action classes, with each sample represented in two modalities: raw video frames and 2D skeletal keypoints. We benchmark several state-of-the-art video action recognition and skeleton-based action classification methods to establish performance baselines for this domain and facilitate future research in automated stroke rehabilitation assessment.

📄 PDF Abstract BibTeX arXiv:2509.07994

Code (0)

등록된 구현이 없습니다.

Tasks

Action ClassificationAction Recognition

Similar Papers 제목 키워드 기반

HAVEN: Hierarchically Aligned Multimodal Benchmark for Unified Video Understanding

2026-05-19 · Mengqi Shi, Haopeng Zhang arxiv

While Multimodal Large Language Models (MLLMs) exhibit strong performance on standard video tasks, their ability to faithfully summarize and reason over complex narratives remains poorly evaluated. Existing summarization…

GEM: A General Evaluation Benchmark for Multimodal Tasks

2021-06-18 · Findings (ACL) 2021 8 · Lin Su, Nan Duan, Edward Cui, Lei Ji 외

In this paper, we present GEM as a General Evaluation benchmark for Multimodal tasks. Different from existing datasets such as GLUE, SuperGLUE, XGLUE and XTREME that mainly focus on natural language tasks, GEM is a large…

Multimodal Fake News Video Explanation: Dataset, Analysis and Evaluation

2025-01-15 · Lizhi Chen, Zhong Qian, Peifeng Li, Qiaoming Zhu

Multimodal fake news videos are difficult to interpret because they require comprehensive consideration of the correlation and consistency between multiple modes. Existing methods deal with fake news videos as a classifi…

DecoderExplanation GenerationRelationSentence

CFVBench: A Comprehensive Video Benchmark for Fine-grained Multimodal Retrieval-Augmented Generation

2025-10-10 · Kaiwen Wei, Xiao Liu, Jie Zhang, Zijian Wang 외 arxiv

Multimodal Retrieval-Augmented Generation (MRAG) enables Multimodal Large Language Models (MLLMs) to generate responses with external multimodal evidence, and numerous video-based MRAG benchmarks have been proposed to ev…

Scene Understanding

LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

2024-07-22 · HaoNing Wu, Dongxu Li, Bei Chen, Junnan Li

Large multimodal models (LMMs) are processing increasingly longer and richer inputs. Albeit the progress, few public benchmark is available to measure such development. To mitigate this gap, we introduce LongVideoBench, …

Multiple-choiceQuestion AnsweringVideo Question AnsweringVideo Understanding