paper-with-me

홈 › Papers

Catalogue Grounded Multimodal Attribution for Museum Video under Resource and Regulatory Constraints

2026-03-11 · Minsak Nanang, Adrian Hilton, Armin Mustafa arxiv

Audiovisual (AV) archives in museums and galleries are growing rapidly, but much of this material remains effectively locked away because it lacks consistent, searchable metadata. Existing method for archiving requires extensive manual effort. We address this by automating the most labour intensive part of the workflow: catalogue style metadata curation for in gallery video, grounded in an existing collection database. Concretely, we propose catalogue-grounded multimodal attribution for museum AV content using an open, locally deployable video language model. We design a multi pass pipeline that (i) summarises artworks in a video, (ii) generates catalogue style descriptions and genre labels, and (iii) attempts to attribute title and artist via conservative similarity matching to the structured catalogue. Early deployments on a painting catalogue suggest that this framework can improve AV archive discoverability while respecting resource constraints, data sovereignty, and emerging regulation, offering a transferable template for application-driven machine learning in other high-stakes domains.

📄 PDF Abstract BibTeX arXiv:2603.11147

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Handcrafted Feature-Assisted One-Class Learning for Artist Authentication in Historical Drawings

2026-01-13 · Hassan Ugail, Jan Ritch-Frel, Irina Matuzava arxiv

Authentication and attribution of works on paper remain persistent challenges in cultural heritage, particularly when the available reference corpus is small and stylistic cues are primarily expressed through line and li…

MUSEKG: A Knowledge Graph Over Museum Collections

2025-11-20 · Jinhao Li, Jianzhong Qi, Soyeon Caren Han, Eun-Jung Holden arxiv

Digitisation in the cultural heritage sector has produced large but fragmented repositories of museum collection data, spanning structured catalogue records, images, and unstructured descriptions. Existing museum informa…

Answer Generation

TimeLens: On-Device Artifact Recognition with Retrieval-Augmented Question Answering for the Grand Egyptian Museum

2026-06-11 · Rawan Hesham, Ali Ashraf, Amr Ahmed, Malak Alaa 외 arxiv

TimeLens is an AI-powered bilingual mobile guide for the Grand Egyptian Museum (GEM). Pointing a phone at an exhibit, a visitor sees the artifact recognized in real time and can ask follow-up questions answered in Englis…

Question Answering

MyoMechanix: Biomechanically-Grounded Compositional Skilled Activity Understanding and Coaching

2026-08-26 · Hao Yin, Paritosh Parmar, Lijun Gu, Lin Xu 외 arxiv

Existing action quality assessment (AQA) datasets and methods rely primarily on visual inputs such as RGB and pose, overlooking physiological dynamics such as muscle mechanics and often modeling actions as monolithic pat…

Action Quality AssessmentAction Understanding

Multimodal Fact-Level Attribution for Verifiable Reasoning

2026-02-12 · David Wan, Han Wang, Ziyang Wang, Elias Stengel-Eskin 외 arxiv

Multimodal large language models (MLLMs) are increasingly used for real-world tasks involving multi-step reasoning and long-form generation, where reliability requires grounding model outputs in heterogeneous input sourc…

Multimodal Reasoning