paper-with-me

홈 › Papers

Make VLM Recognize Visual Hallucination on Cartoon Character Image with Pose Information

2024-03-22 · Bumsoo Kim, Wonseop Shin, Kyuchul Lee, Yonghoon Jung, Sanghyun Seo

Leveraging large-scale Text-to-Image (TTI) models have become a common technique for generating exemplar or training dataset in the fields of image synthesis, video editing, 3D reconstruction. However, semantic structural visual hallucinations involving perceptually severe defects remain a concern, especially in the domain of non-photorealistic rendering (NPR) such as cartoons and pixelization-style character. To detect these hallucinations in NPR, We propose a novel semantic structural hallucination detection system using Vision-Language Model (VLM). Our approach is to leverage the emerging capability of large language model, in-context learning which denotes that VLM has seen some examples by user for specific downstream task, here hallucination detection. Based on in-context learning, we introduce pose-aware in-context visual learning (PA-ICVL) which improve the overall performance of VLM by further inputting visual data beyond prompts, RGB images and pose information. By incorporating pose guidance, we enable VLMs to make more accurate decisions. Experimental results demonstrate significant improvements in identifying visual hallucinations compared to baseline methods relying solely on RGB images. Within selected two VLMs, GPT-4v, Gemini pro vision, our proposed PA-ICVL improves the hallucination detection with 50% to 78%, 57% to 80%, respectively. This research advances a capability of TTI models toward real-world applications by mitigating visual hallucinations via in-context visual learning, expanding their potential in non-photorealistic domains. In addition, it showcase how users can boost the downstream-specialized capability of open VLM by harnessing additional conditions. We collect synthetic cartoon-hallucination dataset with TTI models, this dataset and final tuned VLM will be publicly available.

📄 PDF Abstract BibTeX arXiv:2403.15048

Code (0)

등록된 구현이 없습니다.

Tasks

3D ReconstructionHallucinationImage GenerationIn-Context LearningLanguage ModelingLanguage ModellingLarge Language ModelVideo Editing

Similar Papers 제목 키워드 기반

Instance-guided Cartoon Editing with a Large-scale Dataset

2023-12-04 · Jian Lin, Chengze Li, Xueting Liu, Zhongping Ge

Cartoon editing, appreciated by both professional illustrators and hobbyists, allows extensive creative freedom and the development of original narratives within the cartoon domain. However, the existing literature on ca…

Image SegmentationSegmentationSemantic Segmentation

CogCartoon: Towards Practical Story Visualization

2023-12-17 · Zhongyang Zhu, Jie Tang

The state-of-the-art methods for story visualization demonstrate a significant demand for training data and storage, as well as limited flexibility in story presentation, thereby rendering them impractical for real-world…

Story Visualization

RaBit: Parametric Modeling of 3D Biped Cartoon Characters with a Topological-consistent Dataset

2023-03-22 · CVPR 2023 1 · Zhongjin Luo, Shengcai Cai, Jinguo Dong, Ruibo Ming 외

Assisting people in efficiently producing visually plausible 3D characters has always been a fundamental research topic in computer vision and computer graphics. Recent learning-based approaches have achieved unprecedent…

Face to Cartoon Incremental Super-Resolution using Knowledge Distillation

2024-01-27 · Trinetra Devkatte, Shiv Ram Dubey, Satish Kumar Singh, Abdenour Hadid

Facial super-resolution/hallucination is an important area of research that seeks to enhance low-resolution facial images for a variety of applications. While Generative Adversarial Networks (GANs) have shown promise in …

HallucinationIncremental LearningKnowledge DistillationSuper-Resolution

Make-It-Vivid: Dressing Your Animatable Biped Cartoon Characters from Text

2024-03-25 · CVPR 2024 1 · Junshu Tang, Yanhong Zeng, Ke Fan, Xuheng Wang 외

Creating and animating 3D biped cartoon characters is crucial and valuable in various applications. Compared with geometry, the diverse texture design plays an important role in making 3D biped cartoon characters vivid a…

Question AnsweringTexture Synthesis