paper-with-me

홈 › Papers

From 3D Pose to Prose: Biomechanics-Grounded Vision--Language Coaching

2026-03-27 · Yuyang Ji, Yixuan Shen, Shengjie Zhu, Yu Kong, Feng Liu arxiv

We present BioCoach, a biomechanics-grounded vision--language framework for fitness coaching from streaming video. BioCoach fuses visual appearance and 3D skeletal kinematics, through a novel three-stage pipeline: an exercise-specific degree-of-freedom selector that focuses analysis on salient joints; a structured biomechanical context that pairs individualized morphometrics with cycle and constraint analysis; and a vision--biomechanics conditioned feedback module that applies cross-attention to generate precise, actionable text. Using parameter-efficient training that freezes the vision and language backbones, BioCoach yields transparent, personalized reasoning rather than pattern matching. To enable learning and fair evaluation, we augment QEVD-fit-coach with biomechanics-oriented feedback to create QEVD-bio-fit-coach, and we introduce a biomechanics-aware LLM judge metric. BioCoach delivers clear gains on QEVD-bio-fit-coach across lexical and judgment metrics while maintaining temporal triggering; on the original QEVD-fit-coach, it improves text quality and correctness with near-parity timing, demonstrating that explicit kinematics and constraints are key to accurate, phase-aware coaching.

📄 PDF Abstract BibTeX arXiv:2603.26938

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

From Pixels to Prose: A Large Dataset of Dense Image Captions

2024-06-14 · Vasu Singla, Kaiyu Yue, Sukriti Paul, Reza Shirkavand 외

Training large vision-language models requires extensive, high-quality image-text pairs. Existing web-scraped datasets, however, are noisy and lack detailed image descriptions. To bridge this gap, we introduce PixelProse…

Image Captioning

SIMSPINE: A Biomechanics-Aware Simulation Framework for 3D Spine Motion Annotation and Benchmarking

2026-02-24 · Muhammad Saif Ullah Khan, Didier Stricker arxiv

Modeling spinal motion is fundamental to understanding human biomechanics, yet remains underexplored in computer vision due to the spine's complex multi-joint kinematics and the lack of large-scale 3D annotations. We pre…

Does ISO-Grounded NFR Specification Improve LLM Code Generation? A Comparison of Rich and Structured Interventions against a Natural-Language Baseline

2026-08-13 · Joào Pedro Monteiro Pereira, Vinicius Cardoso Garcia arxiv

In LLM-based code generation, Non-Functional Requirements (NFRs) are often specified as terse one-line phrases. We ask whether grounding those specifications in ISO/IEC 25010 Quality Model, either as rich natural-languag…

Code Generation

EditSleuth: A Dataset of Grounded Reasoning Chains for Image-Edit Forensics

2026-05-09 · Van-Loc Nguyen, AprilPyone MaungMaung, Minh-Triet Tran, Isao Echizen arxiv

Forensic analysis of AI-edited images requires more than binary real-versus-fake prediction: a useful system should localize the edit, identify its semantic type, and ground its decisions in visual evidence. Existing ima…

Image Manipulation

Translating Classical Poetry into Modern Prose

2026-06-01 · Chalamalasetti Kranti, Sowmya Vajjala arxiv

We introduce Padyam2Gadyam a dataset for the task of poem-to-prose translation from 13th-17th Century Telugu Classical Poetry to contemporary Telugu and English prose. The dataset consists of 600 poems and their human-ve…