paper-with-me

Papers

SimplerVoice: A Key Message & Visual Description Generator System for Illiteracy

2018-11-03 · Minh N. B. Nguyen, Samuel Thomas, Anne E. Gattiker, Sujatha Kashyap, Kush R. Varshney

We introduce SimplerVoice: a key message and visual description generator system to help low-literate adults navigate the information-dense world with confidence, on their own. SimplerVoice can automatically generate sensible sentences describing an unknown object, extract semantic meanings of the object usage in the form of a query string, then, represent the string as multiple types of visual guidance (pictures, pictographs, etc.). We demonstrate SimplerVoice system in a case study of generating grocery products' manuals through a mobile application. To evaluate, we conducted a user study on SimplerVoice's generated description in comparison to the information interpreted by users from other methods: the original product package and search engines' top result, in which SimplerVoice achieved the highest performance score: 4.82 on 5-point mean opinion score scale. Our result shows that SimplerVoice is able to provide low-literate end-users with simple yet informative components to help them understand how to use the grocery products, and that the system may potentially provide benefits in other real-world use cases

📄 PDF Abstract BibTeX arXiv:1811.01299

Code (0)

등록된 구현이 없습니다.

Tasks

Navigate

Similar Papers 제목 키워드 기반

Message Passing Multi-Agent GANs

2016-12-05 · Arnab Ghosh, Viveka Kulharia, Vinay Namboodiri

Communicating and sharing intelligence among agents is an important facet of achieving Artificial General Intelligence. As a first step towards this challenge, we introduce a novel framework for image generation: Message…

Image Generation

NavHint: Vision and Language Navigation Agent with a Hint Generator

2024-02-04 · Yue Zhang, Quan Guo, Parisa Kordjamshidi

Existing work on vision and language navigation mainly relies on navigation-related losses to establish the connection between vision and language modalities, neglecting aspects of helping the navigation agent build a de…

Vision and Language Navigation

LLM4VG: Large Language Models Evaluation for Video Grounding

2023-12-21 · Wei Feng, Xin Wang, Hong Chen, Zeyang Zhang 외

Recently, researchers have attempted to investigate the capability of LLMs in handling videos and proposed several video LLM models. However, the ability of LLMs to handle video grounding (VG), which is an important time…

Image CaptioningVideo GroundingVisual Question Answering (VQA)

Reranking Laws for Language Generation: A Communication-Theoretic Perspective

2024-09-11 · António Farinhas, Haau-Sing Li, André F. T. Martins

To ensure large language models (LLMs) are used safely, one must reduce their propensity to hallucinate or to generate unacceptable answers. A simple and often used strategy is to first let the LLM generate multiple hypo…

Code GenerationMachine TranslationRerankingText Generation+1

Keep it Consistent: Topic-Aware Storytelling from an Image Stream via Iterative Multi-agent Communication

2019-11-11 · COLING 2020 8 · Ruize Wang, Zhongyu Wei, Ying Cheng, Piji Li 외

Visual storytelling aims to generate a narrative paragraph from a sequence of images automatically. Existing approaches construct text description independently for each image and roughly concatenate them as a story, whi…

Image CaptioningQuestion GenerationVisual Storytelling