paper-with-me

홈 › Papers

Interactive Attention AI to translate low light photos to captions for night scene understanding in women safety

2022-01-04 · Rajagopal A, Nirmala V, Arun Muthuraj Vedamanickam

There is amazing progress in Deep Learning based models for Image captioning and Low Light image enhancement. For the first time in literature, this paper develops a Deep Learning model that translates night scenes to sentences, opening new possibilities for AI applications in the safety of visually impaired women. Inspired by Image Captioning and Visual Question Answering, a novel Interactive Image Captioning is developed. A user can make the AI focus on any chosen person of interest by influencing the attention scoring. Attention context vectors are computed from CNN feature vectors and user-provided start word. The Encoder-Attention-Decoder neural network learns to produce captions from low brightness images. This paper demonstrates how women safety can be enabled by researching a novel AI capability in the Interactive Vision-Language model for perception of the environment in the night.

📄 PDF Abstract BibTeX arXiv:2201.00969

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderDeep LearningImage CaptioningImage EnhancementLanguage ModelingLanguage ModellingLow-Light Image EnhancementQuestion AnsweringScene UnderstandingVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Event Recognition with Automatic Album Detection based on Sequential Processing, Neural Attention and Image Captioning

2019-11-25 · Andrey V. Savchenko

In this paper a new formulation of event recognition task is examined: it is required to predict event categories in a gallery of images, for which albums (groups of photos corresponding to a single event) are unknown. W…

ClusteringImage Captioning

Fast Preprocessing for Robust Face Sketch Synthesis

2017-08-01 · Yibing Song, Jiawei Zhang, Linchao Bao, Qingxiong Yang

Exemplar-based face sketch synthesis methods usually meet the challenging problem that input photos are captured in different lighting conditions from training photos. The critical step causing the failure is the search …

Face Sketch Synthesis

Aesthetic Critiques Generation for Photos

2017-10-01 · ICCV 2017 10 · Kuang-Yu Chang, Kung-Hung Lu, Chu-Song Chen

It is said that a picture is worth a thousand words. Thus, there are various ways to describe an image, especially in aesthetic quality analysis. Although aesthetic quality assessment has generated a great deal of intere…

Image Captioning

Brazilian Portuguese Image Captioning with Transformers: A Study on Cross-Native-Translated Dataset

2026-01-30 · Gabriel Bromonschenkel, Alessandro L. Koerich, Thiago M. Paixão, Hilário Tomaz Alves de Oliveira arxiv

Image captioning (IC) refers to the automatic generation of natural language descriptions for images, with applications ranging from social media content generation to assisting individuals with visual impairments. While…

Image Captioning

Towers of Babel: Combining Images, Language, and 3D Geometry for Learning Multimodal Vision

2021-08-12 · ICCV 2021 10 · Xiaoshi Wu, Hadar Averbuch-Elor, Jin Sun, Noah Snavely

The abundance and richness of Internet photos of landmarks and cities has led to significant progress in 3D vision over the past two decades, including automated 3D reconstructions of the world's landmarks from tourist p…

3D geometryDescriptiveImage CaptioningMultimodal Reasoning