paper-with-me

홈 › Papers

Low to High Dimensional Modality Hallucination using Aggregated Fields of View

2020-07-13 · Kausic Gunasekar, Qiang Qiu, Yezhou Yang

Real-world robotics systems deal with data from a multitude of modalities, especially for tasks such as navigation and recognition. The performance of those systems can drastically degrade when one or more modalities become inaccessible, due to factors such as sensors' malfunctions or adverse environments. Here, we argue modality hallucination as one effective way to ensure consistent modality availability and thereby reduce unfavorable consequences. While hallucinating data from a modality with richer information, e.g., RGB to depth, has been researched extensively, we investigate the more challenging low-to-high modality hallucination with interesting use cases in robotics and autonomous systems. We present a novel hallucination architecture that aggregates information from multiple fields of view of the local neighborhood to recover the lost information from the extant modality. The process is implemented by capturing a non-linear mapping between the data modalities and the learned mapping is used to aid the extant modality to mitigate the risk posed to the system in the adverse scenarios which involve modality loss. We also conduct extensive classification and segmentation experiments on UWRGBD and NYUD datasets and demonstrate that hallucination allays the negative effects of the modality loss. Implementation and models: https://github.com/kausic94/Hallucination

📄 PDF Abstract BibTeX arXiv:2007.06166

Code (1)

kausic94/Hallucination 공식 구현 tf

Tasks

HallucinationVocal Bursts Intensity Prediction

Similar Papers 제목 키워드 기반

Learning Signal-Agnostic Manifolds of Neural Fields

2021-11-11 · NeurIPS 2021 12 · Yilun Du, Katherine M. Collins, Joshua B. Tenenbaum, Vincent Sitzmann

Deep neural networks have been used widely to learn the latent structure of datasets, across modalities such as images, shapes, and audio signals. However, existing models are generally modality-dependent, requiring cust…

A Survey of Hallucination in Large Visual Language Models

2024-10-20 · Wei Lan, WenYi Chen, Qingfeng Chen, Shirui Pan 외

The Large Visual Language Models (LVLMs) enhances user interaction and enriches user experience by integrating visual modality on the basis of the Large Language Models (LLMs). It has demonstrated their powerful informat…

HallucinationHallucination EvaluationSurvey

HRTransNet: HRFormer-Driven Two-Modality Salient Object Detection

2023-01-08 · Bin Tang, Zhengyi Liu, Yacheng Tan, Qian He

The High-Resolution Transformer (HRFormer) can maintain high-resolution representation and share global receptive fields. It is friendly towards salient object detection (SOD) in which the input and output have the same …

global-optimizationObjectobject-detectionObject Detection+2

Robust Multimodal Large Language Models Against Modality Conflict

2025-07-09 · Zongmeng Zhang, Wengang Zhou, Jie Zhao, Houqiang Li arxiv

Despite the impressive capabilities of multimodal large language models (MLLMs) in vision-language tasks, they are prone to hallucinations in real-world scenarios. This paper investigates the hallucination phenomenon in …

Reinforcement LearningPrompt Engineering

MoD-DPO: Towards Mitigating Cross-modal Hallucinations in Omni LLMs using Modality Decoupled Preference Optimization

2026-03-03 · Ashutosh Chaubey, Jiacheng Pang, Mohammad Soleymani arxiv

Omni-modal large language models (omni LLMs) have recently achieved strong performance across audiovisual understanding tasks, yet they remain highly susceptible to cross-modal hallucinations arising from spurious correl…