paper-with-me

Papers

Rethinking Annotator Simulation: Realistic Evaluation of Whole-Body PET Lesion Interactive Segmentation Methods

2024-04-02 · Zdravko Marinov, Moon Kim, Jens Kleesiek, Rainer Stiefelhagen

Interactive segmentation plays a crucial role in accelerating the annotation, particularly in domains requiring specialized expertise such as nuclear medicine. For example, annotating lesions in whole-body Positron Emission Tomography (PET) images can require over an hour per volume. While previous works evaluate interactive segmentation models through either real user studies or simulated annotators, both approaches present challenges. Real user studies are expensive and often limited in scale, while simulated annotators, also known as robot users, tend to overestimate model performance due to their idealized nature. To address these limitations, we introduce four evaluation metrics that quantify the user shift between real and simulated annotators. In an initial user study involving four annotators, we assess existing robot users using our proposed metrics and find that robot users significantly deviate in performance and annotation behavior compared to real annotators. Based on these findings, we propose a more realistic robot user that reduces the user shift by incorporating human factors such as click variation and inter-annotator disagreement. We validate our robot user in a second user study, involving four other annotators, and show it consistently reduces the simulated-to-real user shift compared to traditional robot users. By employing our robot user, we can conduct more large-scale and cost-efficient evaluations of interactive segmentation models, while preserving the fidelity of real user studies. Our implementation is based on MONAI Label and will be made publicly available.

📄 PDF Abstract BibTeX arXiv:2404.01816

Code (0)

등록된 구현이 없습니다.

Tasks

Interactive SegmentationSegmentation

Similar Papers 제목 키워드 기반

Rethinking the Agreement in Human Evaluation Tasks

2018-08-01 · COLING 2018 8 · Jacopo Amidei, Paul Piwek, Alistair Willis

Human evaluations are broadly thought to be more valuable the higher the inter-annotator agreement. In this paper we examine this idea. We will describe our experiments and analysis within the area of Automatic Question …

Dialogue GenerationQuestion GenerationQuestion-GenerationText Generation

Rethinking the Evaluation and Optimization of LLM-Based Social Simulation

2026-08-20 · Pei Wang, Xu Chen, Ji-Rong Wen arxiv

LLM-based social simulation is a promising complement to traditional methods such as surveys and behavioral experiments. A core question is how to evaluate the fidelity of LLM-simulated human behavior and optimize LLMs t…

Rethinking Ground Truth: A Case Study on Human Label Variation in MLLM Benchmarking

2026-03-20 · Tomas Ruiz, Tanalp Agustoslu, Carsten Schwemmer arxiv

Human Label Variation (HLV), i.e. systematic differences among annotators' judgments, remains underexplored in benchmarks despite rapid progress in large language model (LLM) development. We address this gap by introduci…

Rethinking of Radar's Role: A Camera-Radar Dataset and Systematic Annotator via Coordinate Alignment

2021-05-11 · Yizhou Wang, Gaoang Wang, Hung-Min Hsu, Hui Liu 외

Radar has long been a common sensor on autonomous vehicles for obstacle ranging and speed estimation. However, as a robust sensor to all-weather conditions, radar's capability has not been well-exploited, compared with c…

Autonomous Vehiclesobject-detectionObject DetectionRadar Object Detection

Rethinking Crowd Sourcing for Semantic Similarity

2021-09-24 · Shaul Solomon, Adam Cohn, Hernan Rosenblum, Chezi Hershkovitz 외

Estimation of semantic similarity is crucial for a variety of natural language processing (NLP) tasks. In the absence of a general theory of semantic information, many papers rely on human annotators as the source of gro…

Semantic SimilaritySemantic Textual Similarity