paper-with-me

홈 › Papers

LLMs Can Check Their Own Results to Mitigate Hallucinations in Traffic Understanding Tasks

2024-09-19 · Malsha Ashani Mahawatta Dona, Beatriz Cabrero-Daniel, Yinan Yu, Christian Berger

Today's Large Language Models (LLMs) have showcased exemplary capabilities, ranging from simple text generation to advanced image processing. Such models are currently being explored for in-vehicle services such as supporting perception tasks in Advanced Driver Assistance Systems (ADAS) or Autonomous Driving (AD) systems, given the LLMs' capabilities to process multi-modal data. However, LLMs often generate nonsensical or unfaithful information, known as ``hallucinations'': a notable issue that needs to be mitigated. In this paper, we systematically explore the adoption of SelfCheckGPT to spot hallucinations by three state-of-the-art LLMs (GPT-4o, LLaVA, and Llama3) when analysing visual automotive data from two sources: Waymo Open Dataset, from the US, and PREPER CITY dataset, from Sweden. Our results show that GPT-4o is better at generating faithful image captions than LLaVA, whereas the former demonstrated leniency in mislabeling non-hallucinated content as hallucinations compared to the latter. Furthermore, the analysis of the performance metrics revealed that the dataset type (Waymo or PREPER CITY) did not significantly affect the quality of the captions or the effectiveness of hallucination detection. However, the models showed better performance rates over images captured during daytime, compared to during dawn, dusk or night. Overall, the results show that SelfCheckGPT and its adaptation can be used to filter hallucinations in generated traffic-related image captions for state-of-the-art LLMs.

📄 PDF Abstract BibTeX arXiv:2409.12580

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingHallucinationImage CaptioningText Generation

Similar Papers 제목 키워드 기반

HalluciDoctor: Mitigating Hallucinatory Toxicity in Visual Instruction Data

2023-11-22 · CVPR 2024 1 · Qifan Yu, Juncheng Li, Longhui Wei, Liang Pang 외

Multi-modal Large Language Models (MLLMs) tuned on machine-generated instruction-following data have demonstrated remarkable performance in various multi-modal understanding and generation tasks. However, the hallucinati…

AttributecounterfactualHallucinationHallucination Evaluation+1

Interactive DualChecker for Mitigating Hallucinations in Distilling Large Language Models

2024-08-22 · Meiyun Wang, Masahiro Suzuki, Hiroki Sakaji, Kiyoshi Izumi

Large Language Models (LLMs) have demonstrated exceptional capabilities across various machine learning (ML) tasks. Given the high costs of creating annotated datasets for supervised learning, LLMs offer a valuable alter…

In-Context LearningKnowledge Distillationtoken-classificationToken Classification+1

A Stitch in Time Saves Nine: Detecting and Mitigating Hallucinations of LLMs by Validating Low-Confidence Generation

2023-07-08 · Neeraj Varshney, Wenlin Yao, Hongming Zhang, Jianshu Chen 외

Recently developed large language models have achieved remarkable success in generating fluent and coherent text. However, these models often tend to 'hallucinate' which critically hampers their reliability. In this work…

Hallucination

LLMs Will Always Hallucinate, and We Need to Live With This

2024-09-09 · Sourav Banerjee, Ayushi Agarwal, Saloni Singla

As Large Language Models become more ubiquitous across domains, it becomes important to examine their inherent limitations critically. This work argues that hallucinations in language models are not just occasional error…

Fact CheckingHallucinationintent-classificationIntent Classification+1

KnowHalu: Hallucination Detection via Multi-Form Knowledge Based Factual Checking

2024-04-03 · Jiawei Zhang, Chejian Xu, Yu Gai, Freddy Lecue 외

This paper introduces KnowHalu, a novel approach for detecting hallucinations in text generated by large language models (LLMs), utilizing step-wise reasoning, multi-formulation query, multi-form knowledge for factual ch…

Fact CheckingFormHallucinationRetrieval