paper-with-me

홈 › Papers

Contextual Knowledge Pursuit for Faithful Visual Synthesis

2023-11-29 · Jinqi Luo, Kwan Ho Ryan Chan, Dimitris Dimos, René Vidal

Modern text-to-vision generative models often hallucinate when the prompt describing the scene to be generated is underspecified. In large language models (LLMs), a prevalent strategy to reduce hallucinations is to retrieve factual knowledge from an external database. While such retrieval augmentation strategies have great potential to enhance text-to-vision generators, existing static top-K retrieval methods explore the knowledge pool once, missing the broader context necessary for high-quality generation. Furthermore, LLMs internally possess rich world knowledge learned during large-scale training (parametric knowledge) that could mitigate the need for external data retrieval. This paper proposes Contextual Knowledge Pursuit (CKPT), a framework that leverages the complementary strengths of external and parametric knowledge to help generators produce reliable visual content. Instead of the one-time retrieval of facts from an external database to improve a given prompt, CKPT uses (1) an LLM to decide whether to seek external knowledge or to self-elicit descriptions from LLM parametric knowledge, (2) a knowledge pursuit process to contextually seek and sequentially gather most relevant facts, (3) a knowledge aggregator for prompt enhancement with the gathered fact context, and (4) a filtered fine-tuning objective to improve visual synthesis with richer prompts. We evaluate CKPT across multiple text-driven generative tasks (image, 3D rendering, and video) on datasets of rare objects and daily scenarios. Our results show that CKPT is capable of generating faithful and semantically rich content across diverse visual domains, offering a promising data source for zero-shot synthesis and filtered fine-tuning of text-to-vision generative models.

📄 PDF Abstract BibTeX arXiv:2311.17898

Code (1)

peterljq/contextual-knowledge-pursuit 공식 구현

Tasks

Language ModellingRetrievalWorld Knowledge

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Cross-Document Topic-Aligned Chunking for Retrieval-Augmented Generation

2025-11-08 · Mile Stankovic arxiv

Chunking quality determines RAG system performance. Current methods partition documents individually, but complex queries need information scattered across multiple sources: the knowledge fragmentation problem. We introd…

ContextFocus: Activation Steering for Contextual Faithfulness in Large Language Models

2026-01-07 · Nikhil Anand, Shwetha Somasundaram, Anirudh Phukan, Apoorv Saxena 외 arxiv

Large Language Models (LLMs) encode vast amounts of parametric knowledge during pre-training. As world knowledge evolves, effective deployment increasingly depends on their ability to faithfully follow externally retriev…

Novel-View Acoustic Synthesis

2023-01-20 · CVPR 2023 1 · Changan Chen, Alexander Richard, Roman Shapovalov, Vamsi Krishna Ithapu 외

We introduce the novel-view acoustic synthesis (NVAS) task: given the sight and sound observed at a source viewpoint, can we synthesize the sound of that scene from an unseen target viewpoint? We propose a neural renderi…

Neural RenderingNovel View Synthesis

Investigating the Interplay between Contextual and Parametric Chain-of-Thought Faithfulness under Optimization

2026-05-24 · Jingyi Sun, Qianli Wang, Pepa Atanasova, Nils Feldhus 외 arxiv

Chain-of-Thought (CoT) faithfulness, i.e., whether CoTs genuinely reflect large language models' (LLM) underlying behavior, is typically evaluated with metrics under two disjoint paradigms: contextual faithfulness, measu…

Context-faithful Prompting for Large Language Models

2023-03-20 · Wenxuan Zhou, Sheng Zhang, Hoifung Poon, Muhao Chen

Large language models (LLMs) encode parametric knowledge about world facts and have shown remarkable performance in knowledge-driven NLP tasks. However, their reliance on parametric knowledge may cause them to overlook c…

counterfactualMachine Reading ComprehensionReading ComprehensionRelation Extraction