paper-with-me

홈 › Papers

Towards Real Time Egocentric Segment Captioning for The Blind and Visually Impaired in RGB-D Theatre Images

2023-08-26 · Khadidja Delloul, Slimane Larabi

In recent years, image captioning and segmentation have emerged as crucial tasks in computer vision, with applications ranging from autonomous driving to content analysis. Although multiple solutions have emerged to help blind and visually impaired people move around their environment, few are applications that help them understand and rebuild a scene in their minds through text. Most built models focus on helping users move and avoid obstacles, restricting the number of environments blind and visually impaired people can be in. In this paper, we will propose an approach that helps them understand their surroundings using image captioning. The particularity of our research is that we offer them descriptions with positions of regions and objects regarding them (left, right, front), as well as positional relationships between regions, while we aim to give them access to theatre plays by applying the solution to our TS-RGBD dataset.

📄 PDF Abstract BibTeX arXiv:2308.13892

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingImage Captioning

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Exo2EgoDVC: Dense Video Captioning of Egocentric Procedural Activities Using Web Instructional Videos

2023-11-28 · Takehiko Ohkawa, Takuma Yagi, Taichi Nishimura, Ryosuke Furuta 외

We propose a novel benchmark for cross-view knowledge transfer of dense video captioning, adapting models from web instructional videos with exocentric views to an egocentric view. While dense video captioning (predictin…

Dense Video CaptioningTransfer LearningVideo Captioning

Egocentric Image Captioning for Privacy-Preserved Passive Dietary Intake Monitoring

2021-07-01 · Jianing Qiu, Frank P. -W. Lo, Xiao Gu, Modou L. Jobarteh 외

Camera-based passive dietary intake monitoring is able to continuously capture the eating episodes of a subject, recording rich visual information, such as the type and volume of food being consumed, as well as the eatin…

Food RecognitionImage CaptioningScene Understanding

EgoBlind: Towards Egocentric Visual Assistance for the Blind

2025-03-11 · Junbin Xiao, Nanxin Huang, Hao Qiu, Zhulin Tao 외

We present EgoBlind, the first egocentric VideoQA dataset collected from blind individuals to evaluate the assistive capabilities of contemporary multimodal large language models (MLLMs). EgoBlind comprises 1,392 videos …

HCQA @ Ego4D EgoSchema Challenge 2024

2024-06-22 · Haoyu Zhang, Yuquan Xie, Yisen Feng, Zaijing Li 외

In this report, we present our champion solution for Ego4D EgoSchema Challenge in CVPR 2024. To deeply integrate the powerful egocentric captioning model and question reasoning model, we propose a novel Hierarchical Comp…

Caption GenerationEgoSchemaMultiple-choice+2

It's Just Another Day: Unique Video Captioning by Discriminative Prompting

2024-10-15 · Toby Perrett, Tengda Han, Dima Damen, Andrew Zisserman

Long videos contain many repeating actions, events and shots. These repetitions are frequently given identical captions, which makes it difficult to retrieve the exact desired clip using a text search. In this paper, we …

Video Captioning