paper-with-me

홈 › Papers

Human-AI Collaborative Multi-modal Multi-rater Learning for Endometriosis Diagnosis

2024-09-03 · Hu Wang, David Butler, Yuan Zhang, Jodie Avery, Steven Knox, Congbo Ma, Louise Hull, Gustavo Carneiro

Endometriosis, affecting about 10% of individuals assigned female at birth, is challenging to diagnose and manage. Diagnosis typically involves the identification of various signs of the disease using either laparoscopic surgery or the analysis of T1/T2 MRI images, with the latter being quicker and cheaper but less accurate. A key diagnostic sign of endometriosis is the obliteration of the Pouch of Douglas (POD). However, even experienced clinicians struggle with accurately classifying POD obliteration from MRI images, which complicates the training of reliable AI models. In this paper, we introduce the Human-AI Collaborative Multi-modal Multi-rater Learning (HAICOMM) methodology to address the challenge above. HAICOMM is the first method that explores three important aspects of this problem: 1) multi-rater learning to extract a cleaner label from the multiple "noisy" labels available per training sample; 2) multi-modal learning to leverage the presence of T1/T2 MRI images for training and testing; and 3) human-AI collaboration to build a system that leverages the predictions from clinicians and the AI model to provide more accurate classification than standalone clinicians and AI models. Presenting results on the multi-rater T1/T2 MRI endometriosis dataset that we collected to validate our methodology, the proposed HAICOMM model outperforms an ensemble of clinicians, noisy-label learning models, and multi-rater learning methods.

📄 PDF Abstract BibTeX arXiv:2409.02046

Code (0)

등록된 구현이 없습니다.

Tasks

Diagnostic

Similar Papers 제목 키워드 기반

Preference Adaptive and Sequential Text-to-Image Generation

2024-12-10 · Ofir Nabati, Guy Tennenholtz, ChihWei Hsu, MoonKyung Ryu 외

We address the problem of interactive text-to-image (T2I) generation, designing a reinforcement learning (RL) agent which iteratively improves a set of generated images for a user through a sequence of prompt expansions.…

Image GenerationLanguage ModelingLanguage ModellingReinforcement Learning (RL)+2

Through the Judge's Eyes: Inferred Thinking Traces Improve Reliability of LLM Raters

2025-10-29 · Xingjian Zhang, Tianhong Gao, Suliang Jin, Tianhao Wang 외 arxiv

Large language models (LLMs) are increasingly used as raters for evaluation tasks. However, their reliability is often limited for subjective tasks, when human judgments involve subtle reasoning beyond annotation labels.…

Same Words, Different Judgments: How Preferences Vary Across Modalities

2026-02-26 · Aaron Broukhim, Nadir Weibel, Eshin Jolly arxiv

Preference-based reinforcement learning (PbRL) is the dominant framework for aligning AI systems to human preferences. However, evaluation protocols for such data were designed for text and have not been validated for sp…

Reinforcement Learning

Exploring Human-AI Complementarity in CPS Diagnosis Using Unimodal and Multimodal BERT Models

2025-07-19 · Kester Wong, Sahan Bulathwela, Mutlu Cukurova arxiv

Detecting collaborative problem solving (CPS) indicators from dialogue using machine learning techniques is a significant challenge for the field of AI in Education. Recent studies have explored the use of Bidirectional …

Insights on Disagreement Patterns in Multimodal Safety Perception across Diverse Rater Groups

2024-10-22 · Charvi Rastogi, Tian Huey Teh, Pushkar Mishra, Roma Patel 외

AI systems crucially rely on human ratings, but these ratings are often aggregated, obscuring the inherent diversity of perspectives in real-world phenomenon. This is particularly concerning when evaluating the safety of…