Describing image focused in cognitive and visual details for visually impaired people: An approach to generating inclusive paragraphs
Several services for people with visual disabilities have emerged recently due to achievements in Assistive Technologies and Artificial Intelligence areas. Despite the growth in assistive systems availability, there is a lack of services that support specific tasks, such as understanding the image context presented in online content, e.g., webinars. Image captioning techniques and their variants are limited as Assistive Technologies as they do not match the needs of visually impaired people when generating specific descriptions. We propose an approach for generating context of webinar images combining a dense captioning technique with a set of filters, to fit the captions in our domain, and a language model for the abstractive summary task. The results demonstrated that we can produce descriptions with higher interpretability and focused on the relevant information for that group of people by combining image analysis methods and neural language models.
Code (0)
등록된 구현이 없습니다.
Tasks
Dense CaptioningImage CaptioningLanguage ModelingLanguage ModellingSimilar Papers 제목 키워드 기반
Toward Cognitive Supersensing in Multimodal Large Language Model
Multimodal Large Language Models (MLLMs) have achieved remarkable success in open-vocabulary perceptual tasks, yet their ability to solve complex cognitive problems remains limited, especially when visual details are abs…
Visual Question AnsweringReinforcement LearningVisual ReasoningHeatmap-Based Method for Estimating Drivers' Cognitive Distraction
In order to increase road safety, among the visual and manual distractions, modern intelligent vehicles need also to detect cognitive distracted driving (i.e., the drivers mind wandering). In this study, the influence of…
Progressive Prompt Detailing for Improved Alignment in Text-to-Image Generative Models
Text-to-image generative models often struggle with long prompts detailing complex scenes, diverse objects with distinct visual characteristics and spatial relationships. In this work, we propose SCoPE (Scheduled interpo…
Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)Hidden Meanings in Plain Sight: RebusBench for Evaluating Cognitive Visual Reasoning
Large Vision-Language Models (LVLMs) have achieved remarkable proficiency in explicit visual recognition, effectively describing what is directly visible in an image. However, a critical cognitive gap emerges when the vi…
Visual ReasoningFrom Pixels to Feelings: Aligning MLLMs with Human Cognitive Perception of Images
While Multimodal Large Language Models (MLLMs) are adept at answering what is in an image-identifying objects and describing scenes-they often lack the ability to understand how an image feels to a human observer. This g…
Image Generation