paper-with-me

Papers

Enhancing Explainability in Multimodal Large Language Models Using Ontological Context

2024-09-27 · Jihen Amara, Birgitta König-Ries, Sheeba Samuel

Recently, there has been a growing interest in Multimodal Large Language Models (MLLMs) due to their remarkable potential in various tasks integrating different modalities, such as image and text, as well as applications such as image captioning and visual question answering. However, such models still face challenges in accurately captioning and interpreting specific visual concepts and classes, particularly in domain-specific applications. We argue that integrating domain knowledge in the form of an ontology can significantly address these issues. In this work, as a proof of concept, we propose a new framework that combines ontology with MLLMs to classify images of plant diseases. Our method uses concepts about plant diseases from an existing disease ontology to query MLLMs and extract relevant visual concepts from images. Then, we use the reasoning capabilities of the ontology to classify the disease according to the identified concepts. Ensuring that the model accurately uses the concepts describing the disease is crucial in domain-specific applications. By employing an ontology, we can assist in verifying this alignment. Additionally, using the ontology's inference capabilities increases transparency, explainability, and trust in the decision-making process while serving as a judge by checking if the annotations of the concepts by MLLMs are aligned with those in the ontology and displaying the rationales behind their errors. Our framework offers a new direction for synergizing ontologies and MLLMs, supported by an empirical study using different well-known MLLMs.

📄 PDF Abstract BibTeX arXiv:2409.18753

Code (0)

등록된 구현이 없습니다.

Tasks

Image CaptioningQuestion AnsweringVisual Question Answering

Methods 이 논문이 사용한 방법론

Ontology 설명 없음

Similar Papers 제목 키워드 기반

Explainable Multimodal Aspect-Based Sentiment Analysis with Dependency-guided Large Language Model

2026-01-11 · Zhongzheng Wang, Yuanhe Tian, Hongzhi Wang, Yan Song arxiv

Multimodal aspect-based sentiment analysis (MABSA) aims to identify aspect-level sentiments by jointly modeling textual and visual information, which is essential for fine-grained opinion understanding in social media. E…

Sentiment Analysis

An ontological approach to model and query multimodal concurrent linguistic annotations

2012-05-01 · LREC 2012 5 · Julien Seinturier, Elisabeth Murisasco, Emmanuel Bruno, Philippe Blache

This paper focuses on the representation and querying of knowledge-based multimodal data. This work stands in the OTIM project which aims at processing multimodal annotation of a large conversational French speech corpus…

Development of Ontological Knowledge Bases by Leveraging Large Language Models

2026-01-15 · Le Ngoc Luyen, Marie-Hélène Abel, Philippe Gouspillou arxiv

Ontological Knowledge Bases (OKBs) play a vital role in structuring domain-specific knowledge and serve as a foundation for effective knowledge management systems. However, their traditional manual development poses sign…

To the problem of "The Instrumental complex for ontological engineering purpose" software system design

2018-02-10 · A. V. Palagin, N. G. Petrenko, V. Yu. Velychko, K. S. Malakhov 외

The given work describes methodological principles of design instrumental complex of ontological purpose. Instrumental complex intends for the implementation of the integrated information technologies automated build of …

NEURON: A Neuro-symbolic System for Grounded Clinical Explainability

2026-05-02 · Anuradha Chandrasekaran, Dimitrios Zikos, Mutlu Mete, Alan Pang 외 arxiv

Clinical AI adoption is hindered by the black-box/grey-box nature of high-performing models, which lack the ontological grounding and narrative transparency required for professional-level explainability. We present NEUR…

Mortality Prediction