paper-with-me

Papers

A Transformer-based representation-learning model with unified processing of multimodal input for clinical diagnostics

2023-06-01 · Hong-Yu Zhou, Yizhou Yu, Chengdi Wang, Shu Zhang, Yuanxu Gao, Jia Pan, Jun Shao, Guangming Lu, Kang Zhang, Weimin Li

During the diagnostic process, clinicians leverage multimodal information, such as chief complaints, medical images, and laboratory-test results. Deep-learning models for aiding diagnosis have yet to meet this requirement. Here we report a Transformer-based representation-learning model as a clinical diagnostic aid that processes multimodal input in a unified manner. Rather than learning modality-specific features, the model uses embedding layers to convert images and unstructured and structured text into visual tokens and text tokens, and bidirectional blocks with intramodal and intermodal attention to learn a holistic representation of radiographs, the unstructured chief complaint and clinical history, structured clinical information such as laboratory-test results and patient demographic information. The unified model outperformed an image-only model and non-unified multimodal diagnosis models in the identification of pulmonary diseases (by 12% and 9%, respectively) and in the prediction of adverse clinical outcomes in patients with COVID-19 (by 29% and 7%, respectively). Leveraging unified multimodal Transformer-based models may help streamline triage of patients and facilitate the clinical decision process.

📄 PDF Abstract BibTeX arXiv:2306.00864

Code (1)

rl4m/irene 공식 구현 pytorch

Tasks

DiagnosticRepresentation Learning

Similar Papers 제목 키워드 기반

UFO: A UniFied TransfOrmer for Vision-Language Representation Learning

2021-11-19 · JianFeng Wang, Xiaowei Hu, Zhe Gan, Zhengyuan Yang 외

In this paper, we propose a single UniFied transfOrmer (UFO), which is capable of processing either unimodal inputs (e.g., image or language) or multimodal inputs (e.g., the concatenation of the image and the question), …

Image CaptioningImage-text matchingImage-text RetrievalLanguage Modeling+10

PixelBytes: Catching Unified Embedding for Multimodal Generation

2024-09-03 · Fabien Furfaro

This report introduces PixelBytes Embedding, a novel approach for unified multimodal representation learning. Our method captures diverse inputs in a single, cohesive representation, enabling emergent properties for mult…

Mambamultimodal generationRepresentation LearningState Space Models

Meta-Transformer: A Unified Framework for Multimodal Learning

2023-07-20 · Yiyuan Zhang, Kaixiong Gong, Kaipeng Zhang, Hongsheng Li 외

Multimodal learning aims to build models that can process and relate information from multiple modalities. Despite years of development in this field, it still remains challenging to design a unified network for processi…

Time Series

MED-VT++: Unifying Multimodal Learning with a Multiscale Encoder-Decoder Video Transformer

2023-04-12 · CVPR 2023 1 · Rezaul Karim, He Zhao, Richard P. Wildes, Mennatullah Siam

In this paper, we present an end-to-end trainable unified multiscale encoder-decoder transformer that is focused on dense prediction tasks in video. The presented Multiscale Encoder-Decoder Video Transformer (MED-VT) use…

Action SegmentationDecoderOptical Flow EstimationSegmentation+5

UniT: Multimodal Multitask Learning with a Unified Transformer

2021-02-22 · ICCV 2021 10 · Ronghang Hu, Amanpreet Singh

We propose UniT, a Unified Transformer model to simultaneously learn the most prominent tasks across different domains, ranging from object detection to natural language understanding and multimodal reasoning. Based on t…

DecoderMultimodal ReasoningMulti-Task LearningNatural Language Understanding+3