paper-with-me

Papers

EKTVQA: Generalized use of External Knowledge to empower Scene Text in Text-VQA

2021-08-22 · Arka Ujjal Dey, Ernest Valveny, Gaurav Harit

The open-ended question answering task of Text-VQA often requires reading and reasoning about rarely seen or completely unseen scene-text content of an image. We address this zero-shot nature of the problem by proposing the generalized use of external knowledge to augment our understanding of the scene text. We design a framework to extract, validate, and reason with knowledge using a standard multimodal transformer for vision language understanding tasks. Through empirical evidence and qualitative results, we demonstrate how external knowledge can highlight instance-only cues and thus help deal with training data bias, improve answer entity type correctness, and detect multiword named entities. We generate results comparable to the state-of-the-art on three publicly available datasets, under the constraints of similar upstream OCR systems and training data.

📄 PDF Abstract BibTeX arXiv:2108.09717

Code (0)

등록된 구현이 없습니다.

Tasks

Open-Ended Question AnsweringOptical Character Recognition (OCR)Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Retrieval-Augmented Purifier for Robust LLM-Empowered Recommendation

2025-04-03 · Liangbo Ning, Wenqi Fan, Qing Li

Recently, Large Language Model (LLM)-empowered recommender systems have revolutionized personalized recommendation frameworks and attracted extensive attention. Despite the remarkable success, existing LLM-empowered RecS…

Large Language ModelRAGRecommendation SystemsRetrieval+1

KG-ViP: Bridging Knowledge Grounding and Visual Perception in Multi-modal LLMs for Visual Question Answering

2026-01-14 · Zhiyang Li, Ao Ke, Yukun Cao, Xike Xie arxiv

Multi-modal Large Language Models (MLLMs) for Visual Question Answering (VQA) often suffer from dual limitations: knowledge hallucination and insufficient fine-grained visual perception. Crucially, we identify that commo…

Visual Question Answering

Agent-centric learning: from external reward maximization to internal knowledge curation

2025-07-29 · Hanqi Zhou, Fryderyk Mantiuk, David G. Nagy, Charley M. Wu arxiv

The pursuit of general intelligence has traditionally centered on external objectives: an agent's control over its environments or mastery of specific tasks. This external focus, however, can produce specialized agents t…

Occlusion-Free Scene Recovery via Neural Radiance Fields

2023-01-01 · CVPR 2023 1 · Chengxuan Zhu, Renjie Wan, Yunkai Tang, Boxin Shi

Our everyday lives are filled with occlusions that we strive to see through. By aggregating desired background information from different viewpoints, we can easily eliminate such occlusions without any external occlu…

NeRFPosition

Pseudo-Generalized Dynamic View Synthesis from a Video

2023-10-12 · Xiaoming Zhao, Alex Colburn, Fangchang Ma, Miguel Angel Bautista 외

Rendering scenes observed in a monocular video from novel viewpoints is a challenging problem. For static scenes the community has studied both scene-specific optimization techniques, which optimize on every test scene, …

Novel View Synthesis