paper-with-me

Papers

NarrationBot and InfoBot: A Hybrid System for Automated Video Description

2021-11-07 · Shasta Ihorn, Yue-Ting Siu, Aditya Bodi, Lothar Narins, Jose M. Castanon, Yash Kant, Abhishek Das, Ilmi Yoon, Pooyan Fazli

Video accessibility is crucial for blind and low vision users for equitable engagements in education, employment, and entertainment. Despite the availability of professional and amateur services and tools, most human-generated descriptions are expensive and time consuming. Moreover, the rate of human-generated descriptions cannot match the speed of video production. To overcome the increasing gaps in video accessibility, we developed a hybrid system of two tools to 1) automatically generate descriptions for videos and 2) provide answers or additional descriptions in response to user queries on a video. Results from a mixed-methods study with 26 blind and low vision individuals show that our system significantly improved user comprehension and enjoyment of selected videos when both tools were used in tandem. In addition, participants reported no significant difference in their ability to understand videos when presented with autogenerated descriptions versus human-revised autogenerated descriptions. Our results demonstrate user enthusiasm about the developed system and its promise for providing customized access to videos. We discuss the limitations of the current work and provide recommendations for the future development of automated video description tools.

📄 PDF Abstract BibTeX arXiv:2111.03994

Code (0)

등록된 구현이 없습니다.

Tasks

Video Description

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

Towards End-to-End Reinforcement Learning of Dialogue Agents for Information Access

2016-09-03 · ACL 2017 7 · Bhuwan Dhingra, Lihong Li, Xiujun Li, Jianfeng Gao 외

This paper proposes KB-InfoBot -- a multi-turn dialogue agent which helps users search Knowledge Bases (KBs) without composing complicated queries. Such goal-oriented dialogue agents typically need to interact with an ex…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)Retrieval+1

Video Intelligence as a component of a Global Security system

2022-01-12 · Dominique Verdejo, Eunika Mercier-Laurent

This paper describes the evolution of our research from video analytics to a global security system with focus on the video surveillance component. Indeed video surveillance has evolved from a commodity security tool up …

Management

Complex Human Action Recognition in Live Videos Using Hybrid FR-DL Method

2020-07-06 · Fatemeh Serpush, Mahdi Rezaei

Automated human action recognition is one of the most attractive and practical research fields in computer vision, in spite of its high computational costs. In such systems, the human action labelling is based on the app…

Action RecognitionArticlesBenchmarkingfeature selection+1

Video-Based MPAA Rating Prediction: An Attention-Driven Hybrid Architecture Using Contrastive Learning

2025-09-08 · Dipta Neogi, Nourash Azmine Chowdhury, Muhammad Rafsan Kabir, Mohammad Ashrafuzzaman Khan arxiv

The rapid growth of visual content consumption across platforms necessitates automated video classification for age-suitability standards like the MPAA rating system (G, PG, PG-13, R). Traditional methods struggle with l…

Video ClassificationContrastive Learning

From Skeletons to Semantics: Design and Deployment of a Hybrid Edge-Based Action Detection System for Public Safety

2026-03-31 · Ganen Sethupathy, Lalit Dumka, Jan Schagen arxiv

Public spaces such as transport hubs, city centres, and event venues require timely and reliable detection of potentially violent behaviour to support public safety. While automated video analysis has made significant pr…

Action Detection