paper-with-me

홈 › Papers

FERA: A Pose-Based Framework for Rule-Grounded Multimedia Decision Support with a Foil Fencing Case Study

2025-09-23 · Ziwen Chen, Zhong Wang arxiv

Multimedia decision support requires more than recognition; it requires explicit state estimates that can be checked against rules, audited by humans, and consumed by downstream decision logic. We present the FEncing Referee Assistant (FERA), a pose-based framework for this setting, and study it through foil fencing, where decisions depend on fast bilateral motion and right-of-way rules. The framework separates canonical participant tracking, kinematic tokenization, calibrated temporal perception, a compact structured decision layer, and an explanation-oriented retrieval interface. We also release an audited benchmark with adjudicated labels and fixed folds for reproducible evaluation. Under a shared protocol, a lightweight lifted-depth sidecar strengthens the best graph-based perception model, while a compact structured classifier on the fixed two-dimensional token stream reaches 0.624 accuracy and a 0.632 macro-averaged F1 score on the final Left / Right / None decision. The case study supports a broader design lesson: keep the boundary between perception and rule application explicit, preserve uncertainty, and choose the perception front end according to the downstream operating point.

📄 PDF Abstract BibTeX arXiv:2509.18527

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Informative Visual Storytelling with Cross-modal Rules

2019-07-07 · Jiacheng Li, Haizhou Shi, Siliang Tang, Fei Wu 외

Existing methods in the Visual Storytelling field often suffer from the problem of generating general descriptions, while the image contains a lot of meaningful contents remaining unnoticed. The failure of informative st…

DecoderStory GenerationVisual Storytelling

Transferable Persona-Grounded Dialogues via Grounded Minimal Edits

2021-09-16 · EMNLP 2021 11 · Chen Henry Wu, Yinhe Zheng, Xiaoxi Mao, Minlie Huang

Grounded dialogue models generate responses that are grounded on certain concepts. Limited by the distribution of grounded dialogue data, models trained on such data face the transferability challenges in terms of the da…

Increased-confidence adversarial examples for deep learning counter-forensics

2020-05-12 · Wenjie Li, Benedetta Tondi, Rongrong Ni, Mauro Barni

Transferability of adversarial examples is a key issue to apply this kind of attacks against multimedia forensics (MMF) techniques based on Deep Learning (DL) in a real-life setting. Adversarial example transferability, …

Deep LearningImage Forensics

Visual Semantic Multimedia Event Model for Complex Event Detection in Video Streams

2020-09-30 · Piyush Yadav, Edward Curry

Multimedia data is highly expressive and has traditionally been very difficult for a machine to interpret. Middleware systems such as complex event processing (CEP) mine patterns from data streams and send notifications …

Event Detection

MMTB: Evaluating Terminal Agents on Multimedia-File Tasks

2026-05-08 · Chiyeong Heo, Jaechang Kim, Junhyuk Kwon, Hoyoung Kim 외 arxiv

Terminals provide a powerful interface for AI agents by exposing diverse tools for automating complex workflows, yet existing terminal-agent benchmarks largely focus on tasks grounded in text, code, and structured files.…