Learning With Side Information Through Modality Hallucination
We present a modality hallucination architecture for training an RGB object detection model which incorporates depth side information at training time. Our convolutional hallucination network learns a new and complementary RGB image representation which is taught to mimic convolutional mid-level features from a depth network. At test time images are processed jointly through the RGB and hallucination networks to produce improved detection performance. Thus, our method transfers information commonly extracted from depth training data to a network which can extract that information from the RGB counterpart. We present results on the standard NYUDv2 dataset and report improvement on the RGB detection task.
Code (0)
등록된 구현이 없습니다.
Tasks
Hallucinationobject-detectionObject DetectionSimilar Papers 제목 키워드 기반
Low to High Dimensional Modality Hallucination using Aggregated Fields of View
Real-world robotics systems deal with data from a multitude of modalities, especially for tasks such as navigation and recognition. The performance of those systems can drastically degrade when one or more modalities bec…
HallucinationVocal Bursts Intensity PredictionMAD: Modality-Adaptive Decoding for Mitigating Cross-Modal Hallucinations in Multimodal Large Language Models
Multimodal Large Language Models (MLLMs) suffer from cross-modal hallucinations, where one modality inappropriately influences generation about another, leading to fabricated output. This exposes a more fundamental defic…
Multimodal ReasoningCure or Poison? Embedding Instructions Visually Alters Hallucination in Vision-Language Models
Vision-Language Models (VLMs) often suffer from hallucination, partly due to challenges in aligning multimodal information. We propose Prompt-in-Image, a simple method that embeds textual instructions directly into image…
Modality Bias in LVLMs: Analyzing and Mitigating Object Hallucination via Attention Lens
Large vision-language models (LVLMs) have demonstrated remarkable multimodal comprehension and reasoning capabilities, but they still suffer from severe object hallucination. Previous studies primarily attribute the flaw…
Understanding the Role of Hallucination in Reinforcement Post-Training of Multimodal Reasoning Models
The recent success of reinforcement learning (RL) in large reasoning models has inspired the growing adoption of RL for post-training Multimodal Large Language Models (MLLMs) to enhance their visual reasoning capabilitie…
Reinforcement LearningMultimodal ReasoningVisual Reasoning