paper-with-me

Papers

LIMSSR: LLM-Driven Sequence-to-Score Reasoning under Training-Time Incomplete Multimodal Observations

2026-05-01 · Huangbiao Xu, Huanqi Wu, Xiao Ke, Yuxin Peng arxiv

Real-world multimodal learning is often hindered by missing modalities. While Incomplete Multimodal Learning (IML) has gained traction, existing methods typically rely on the unrealistic assumption of full-modal availability during training to provide reconstruction supervision or cross-modal priors. This paper tackles the more challenging setting of IML under training-time incomplete observations, which precludes reliance on a ``God's eye view'' of complete data. We propose LIMSSR (LLM-Driven Incomplete Multimodal Sequence-to-Score Reasoning), a framework that reformulates this challenge as a conditional sequence reasoning task. LIMSSR leverages the semantic reasoning capabilities of Large Language Models via Prompt-Guided Context-Aware Modality Imputation and Multidimensional Representation Fusion to infer latent semantics from available contexts without direct reconstruction. To mitigate hallucinations, we introduce a Mask-Aware Dual-Path Aggregation to dynamically calibrate inference uncertainty. Extensive experiments on three Action Quality Assessment datasets demonstrate that LIMSSR significantly outperforms state-of-the-art baselines without relying on complete training data, establishing a new paradigm for data-efficient multimodal learning. Code is available at https://github.com/XuHuangbiao/LIMSSR.

📄 PDF Abstract BibTeX arXiv:2605.00434

Code (0)

등록된 구현이 없습니다.

Tasks

Action Quality Assessment

Similar Papers 제목 키워드 기반

Quality-Driven Agentic Reasoning for LLM-Assisted Software Design: Questions-of-Thoughts (QoT) as a Time-Series Self-QA Chain

2026-03-10 · Yen-Ku Liu, Yun-Cheng Tsai arxiv

Recent advances in large language models (LLMs) have accelerated AI-assisted software development, yet practical deployment remains constrained by incomplete implementations, weak modularization, and inconsistent securit…

Biological Reasoning-Informed Regression for Interpretable Regulatory DNA Activity Prediction

2026-06-06 · Yi Duan, Zhao Yang, Jiwei Zhu, Ying Ba 외 arxiv

DNA cis-regulatory elements (CREs) such as enhancers control gene expression levels. Accurately predicting regulatory activity from DNA sequences is valuable but challenging, as it requires understanding complex biologic…

Activity Prediction

Where to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video Understanding

2026-04-19 · Shaoguang Wang, Weiyu Guo, Ziyang Chen, Xuming Hu 외 arxiv

Long video understanding remains a formidable challenge for Multimodal Large Language Models (MLLMs) due to the prohibitive computational cost of processing dense frame sequences. Prevailing solutions, which select a key…

Visual Grounding

QTrack: Query-Driven Reasoning for Multi-modal MOT

2026-03-14 · Tajamul Ashraf, Tavaheed Tariq, Sonia Yadav, Abrar Ul Riyaz 외 arxiv

Multi-object tracking (MOT) has traditionally focused on estimating trajectories of all objects in a video, without selectively reasoning about user-specified targets under semantic instructions. In this work, we introdu…

Natural Language QueriesMulti-Object TrackingMultimodal Reasoning

PPCR-IM: A System for Multi-layer DAG-based Public Policy Consequence Reasoning and Social Indicator Mapping

2026-02-25 · Zichen Song, Weijia Li arxiv

Public policy decisions are typically justified using a narrow set of headline indicators, leaving many downstream social impacts unstructured and difficult to compare across policies. We propose PPCR-IM, a system for mu…