paper-with-me

Papers

Visually Dehallucinative Instruction Generation: Know What You Don't Know

2024-02-15 · Sungguk Cha, Jusung Lee, Younghyun Lee, Cheoljong Yang

"When did the emperor Napoleon invented iPhone?" Such hallucination-inducing question is well known challenge in generative language modeling. In this study, we present an innovative concept of visual hallucination, referred to as "I Know (IK)" hallucination, to address scenarios where "I Don't Know" is the desired response. To effectively tackle this issue, we propose the VQAv2-IDK benchmark, the subset of VQAv2 comprising unanswerable image-question pairs as determined by human annotators. Stepping further, we present the visually dehallucinative instruction generation method for IK hallucination and introduce the IDK-Instructions visual instruction database. Our experiments show that current methods struggle with IK hallucination. Yet, our approach effectively reduces these hallucinations, proving its versatility across different frameworks and datasets.

📄 PDF Abstract BibTeX arXiv:2402.09717

Code (1)

ncsoft/idk 공식 구현

Tasks

HallucinationLanguage ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

Visually Dehallucinative Instruction Generation

2024-02-13 · Sungguk Cha, Jusung Lee, Younghyun Lee, Cheoljong Yang

In recent years, synthetic visual instructions by generative language model have demonstrated plausible text generation performance on the visual question-answering tasks. However, challenges persist in the hallucination…

HallucinationLanguage ModelingLanguage ModellingQuestion Answering+2

Vision-and-Language Navigation: Interpreting visually-grounded navigation instructions in real environments

2017-11-20 · CVPR 2018 6 · Peter Anderson, Qi Wu, Damien Teney, Jake Bruce 외

A robot that can carry out a natural-language instruction has been a dream since before the Jetsons cartoon series imagined a life of leisure mediated by a fleet of attentive robot helpers. It is a dream that remains stu…

Reinforcement LearningTranslationVision and Language NavigationVisual Navigation+2

Coherent Zero-Shot Visual Instruction Generation

2024-06-06 · Quynh Phung, Songwei Ge, Jia-Bin Huang

Despite the advances in text-to-image synthesis, particularly with diffusion models, generating visual instructions that require consistent representation and smooth state transitions of objects across sequential steps r…

Image GenerationReading Comprehension

LaF-GRPO: In-Situ Navigation Instruction Generation for the Visually Impaired via GRPO with LLM-as-Follower Reward

2025-06-04 · Yi Zhao, Siqi Wang, Jing Li

Navigation instruction generation for visually impaired (VI) individuals (NIG-VI) is critical yet relatively underexplored. This study, hence, focuses on producing precise, in-situ, step-by-step navigation instructions t…

Language ModelingLanguage Modelling

What You Say Is What You Show: Visual Narration Detection in Instructional Videos

2023-01-05 · Kumar Ashutosh, Rohit Girdhar, Lorenzo Torresani, Kristen Grauman

Narrated ''how-to'' videos have emerged as a promising data source for a wide range of learning problems, from learning visual representations to training robot policies. However, this data is extremely noisy, as the nar…