paper-with-me

Papers

GigaPevt: Multimodal Medical Assistant

2024-02-26 · Pavel Blinov, Konstantin Egorov, Ivan Sviridov, Nikolay Ivanov, Stepan Botman, Evgeniy Tagin, Stepan Kudin, Galina Zubkova, Andrey Savchenko

Building an intelligent and efficient medical assistant is still a challenging AI problem. The major limitation comes from the data modality scarceness, which reduces comprehensive patient perception. This demo paper presents the GigaPevt, the first multimodal medical assistant that combines the dialog capabilities of large language models with specialized medical models. Such an approach shows immediate advantages in dialog quality and metric performance, with a 1.18% accuracy improvement in the question-answering task.

📄 PDF Abstract BibTeX arXiv:2402.16654

Code (0)

등록된 구현이 없습니다.

Tasks

Question Answering

Similar Papers 제목 키워드 기반

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants

2024-12-17 · Hritik Bansal, Daniel Israel, Siyan Zhao, Shufan Li 외

Recent advancements in mixed-modal generative models have enabled flexible integration of information across image-text content. These models have opened new avenues for developing unified biomedical assistants capable o…

Image CaptioningQuestion AnsweringVisual Question Answering

LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day

2023-06-01 · NeurIPS 2023 11 · Chunyuan Li, Cliff Wong, Sheng Zhang, Naoto Usuyama 외

Conversational generative AI has demonstrated remarkable promise for empowering biomedical practitioners, but current investigations focus on unimodal text. Multimodal conversational AI has seen rapid progress by leverag…

Image ClassificationInstruction FollowingLanguage ModellingQuestion Answering+3

OphGLM: Training an Ophthalmology Large Language-and-Vision Assistant based on Instructions and Dialogue

2023-06-21 · Weihao Gao, Zhuo Deng, Zhiyuan Niu, Fuju Rong 외

Large multimodal language models (LMMs) have achieved significant success in general domains. However, due to the significant differences between medical images and text and general web content, the performance of LMMs i…

Instruction FollowingLanguage ModelingLanguage ModellingLarge Language Model+1

LLaVA-Ultra: Large Chinese Language and Vision Assistant for Ultrasound

2024-10-19 · Xuechen Guo, Wenhao Chai, Shi-Yan Li, Gaoang Wang

Multimodal Large Language Model (MLLM) has recently garnered attention as a prominent research focus. By harnessing powerful LLM, it facilitates a transition of conversational generative AI from unimodal text to performi…

Instruction FollowingKnowledge DistillationLanguage ModelingLanguage Modelling+6

Reinforced Correlation Between Vision and Language for Precise Medical AI Assistant

2025-05-06 · Haonan Wang, Jiaji Mao, Lehan Wang, Qixiang Zhang 외

Medical AI assistants support doctors in disease diagnosis, medical image analysis, and report generation. However, they still face significant challenges in clinical use, including limited accuracy with multimodal conte…

Cell SegmentationMedical Image Analysis