paper-with-me

홈 › Papers

Evidence-Grounded Multimodal Knowledge Graph Construction for Multi-Lecture Educational Reasoning

2026-08-04 · Sahil Al Farib, Momota Ahsana Meem, Sheikh Redwanul Islam, Md. Tanvir Raihan arxiv

Lecture videos distribute knowledge across speech, slide text, diagrams, equations, and presentation order, which transcript-only retrieval does not fully preserve. This paper presents an evidence-grounded multimodal pipeline that transcribes lectures, selects semantic anchors, applies optical character recognition (OCR), and uses a vision-language model to extract only concepts and typed relationships supported by transcript, OCR, or visual evidence. Mentions are validated and canonicalized into a provenance-rich knowledge graph. On three neural-network lectures, the pipeline processed 3,118 frames, 756 transcript segments, and 559 anchors. It retained 1,022 concept and 312 relationship mentions, yielding 172 canonical concepts and 282 relationships with 90.38% endpoint coverage. A preliminary three question retrieval test achieved 100% top-1 and top-3 accuracy and 100% mean top-5 recall. The contribution is an auditable construction method rather than a state-of-the-art performance claim.

📄 PDF Abstract BibTeX arXiv:2608.03161

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RetiBridge: Bridging Quantitative Retinal Biomarkers and Qualitative Diagnosis with a Knowledge-Guided Multimodal Large Language Model

2025-10-05 · Zhuangzhi Gao, Hongyi Qin, He Zhao, Qinkai Yu 외 arxiv

Retinal biomarkers captured by color fundus photography and optical coherence tomography provide clinically valuable evidence for both ocular and systemic diseases. Multimodal large language models (MLLMs) have shown pro…

Plan2Map: A Multimodal Benchmark for Document-Grounded Geospatial Boundary Reconstruction from Planning Records

2026-06-01 · Fabian Degen, Oishi Deb, Jindong Gu, Junchi Yu 외 arxiv

Planning records define restrictions over geographic areas, but their source documents often provide only indirect spatial evidence rather than machine-readable boundaries. We introduce Plan2Map, a 208-case multimodal be…

Propagating construction-time knowledge quality into medical question answering: A framework grounded in clinical guidelines

2026-08-28 · Jie Hu, Junjie Wang, Shan Lu, Yifang Hu 외 arxiv

Large language models have facilitated knowledge graph (KG) construction from clinical guidelines, but extracted triples vary in structural validity and evidential support. Meanwhile, graph-augmented question answering (…

Question Answering

CuriosAI Submission to the CASTLE Challenge at EgoVis 2026

2026-05-27 · Yuto Kanda, Hayato Tanoue, Takayuki Hori arxiv

CASTLE 2026 asks 185 multiple-choice questions over 600+ hours of synchronized multi-view egocentric video. We explore two approaches on top of a shared multimodal preprocessing layer, including per-person timelines, spe…

A Multimodal Agentic Pathology Co-pilot via Evidence Grounded Reasoning

2026-06-06 · Zhe Xu, Zhengyu Zhang, Zhiyuan Cai, Jiahao Xu 외 arxiv

Pathology is the cornerstone of modern medicine, where accurate decision-making relies heavily on evidence-based practices. While artificial intelligence (AI) has the potential to transform clinical workflows, the inters…