paper-with-me

홈 › Papers

Helping Visually Impaired People Take Better Quality Pictures

2023-05-14 · Maniratnam Mandal, Deepti Ghadiyaram, Danna Gurari, Alan C. Bovik

Perception-based image analysis technologies can be used to help visually impaired people take better quality pictures by providing automated guidance, thereby empowering them to interact more confidently on social media. The photographs taken by visually impaired users often suffer from one or both of two kinds of quality issues: technical quality (distortions), and semantic quality, such as framing and aesthetic composition. Here we develop tools to help them minimize occurrences of common technical distortions, such as blur, poor exposure, and noise. We do not address the complementary problems of semantic quality, leaving that aspect for future work. The problem of assessing and providing actionable feedback on the technical quality of pictures captured by visually impaired users is hard enough, owing to the severe, commingled distortions that often occur. To advance progress on the problem of analyzing and measuring the technical quality of visually impaired user-generated content (VI-UGC), we built a very large and unique subjective image quality and distortion dataset. This new perceptual resource, which we call the LIVE-Meta VI-UGC Database, contains $40$K real-world distorted VI-UGC images and $40$K patches, on which we recorded $2.7$M human perceptual quality judgments and $2.7$M distortion labels. Using this psychometric resource we also created an automatic blind picture quality and distortion predictor that learns local-to-global spatial quality relationships, achieving state-of-the-art prediction performance on VI-UGC pictures, significantly outperforming existing picture quality models on this unique class of distorted picture data. We also created a prototype feedback system that helps to guide users to mitigate quality issues and take better quality pictures, by creating a multi-task learning framework.

📄 PDF Abstract BibTeX arXiv:2305.08066

Code (1)

mandal-cv/visimpaired 공식 구현 tf

Tasks

Multi-Task Learning

Similar Papers 제목 키워드 기반

Newvision: application for helping blind people using deep learning

2023-11-05 · Kumar Srinivas Bobba, Kartheeban K, Vamsi Krishna Sai Boddu, Vijaya Mani Surendra Bolla 외

As able-bodied people, we often take our vision for granted. For people who are visually impaired, however, their disability can have a significant impact on their daily lives. We are developing proprietary headgear that…

Deep LearningNavigate

Towards Real Time Egocentric Segment Captioning for The Blind and Visually Impaired in RGB-D Theatre Images

2023-08-26 · Khadidja Delloul, Slimane Larabi

In recent years, image captioning and segmentation have emerged as crucial tasks in computer vision, with applications ranging from autonomous driving to content analysis. Although multiple solutions have emerged to help…

Autonomous DrivingImage Captioning

Development of a Neural Network Model for Currency Detection to aid visually impaired people in Nigeria

2025-08-25 · Sochukwuma Nwokoye, Desmond Moru arxiv

Neural networks in assistive technology for visually impaired leverage artificial intelligence's capacity to recognize patterns in complex data. They are used for converting visual data into auditory or tactile represent…

Design of a Mobile Face Recognition System for Visually Impaired Persons

2015-02-03 · Shonal Chaudhry, Rohitash Chandra

It is estimated that 285 million people globally are visually impaired. A majority of these people live in developing countries and are among the elderly population. One of the most difficult tasks faced by the visually …

Face DetectionFace Recognition

Image Captioning as an Assistive Technology: Lessons Learned from VizWiz 2020 Challenge

2020-12-21 · Pierre Dognin, Igor Melnyk, Youssef Mroueh, Inkit Padhi 외

Image captioning has recently demonstrated impressive progress largely owing to the introduction of neural network algorithms trained on curated dataset like MS-COCO. Often work in this field is motivated by the promise …

Image CaptioningNavigate