paper-with-me

Papers

Cleaning and Structuring the Label Space of the iMet Collection 2020

2021-06-01 · Vivien Nguyen, Sunnie S. Y. Kim

The iMet 2020 dataset is a valuable resource in the space of fine-grained art attribution recognition, but we believe it has yet to reach its true potential. We document the unique properties of the dataset and observe that many of the attribute labels are noisy, more than is implied by the dataset description. Oftentimes, there are also semantic relationships between the labels (e.g., identical, mutual exclusion, subsumption, overlap with uncertainty) which we believe are underutilized. We propose an approach to cleaning and structuring the iMet 2020 labels, and discuss the implications and value of doing so. Further, we demonstrate the benefits of our proposed approach through several experiments. Our code and cleaned labels are available at https://github.com/sunniesuhyoung/iMet2020cleaned.

📄 PDF Abstract BibTeX arXiv:2106.00815

Code (1)

sunniesuhyoung/iMet2020cleaned 공식 구현

Tasks

Attribute

Similar Papers 제목 키워드 기반

Data Cleaning Tools for Token Classification Tasks

2021-06-01 · NAACL (DaSH) 2021 6 · Karthik Muthuraman, Frederick Reiss, Hong Xu, Bryan Cutler 외

Human-in-the-loop systems for cleaning NLP training data rely on automated sieves to isolate potentially-incorrect labels for manual review. We have developed a novel technique for flagging potentially-incorrect labels w…

Classificationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+3

MultiMET: A Multimodal Dataset for Metaphor Understanding

2021-08-01 · ACL 2021 5 · Dongyu Zhang, Minghao Zhang, Heting Zhang, Liang Yang 외

Metaphor involves not only a linguistic phenomenon, but also a cognitive phenomenon structuring human thought, which makes understanding it challenging. As a means of cognition, metaphor is rendered by more than texts al…

Inconsistency Ranking-based Noisy Label Detection for High-quality Data

2022-12-01 · Ruibin Yuan, Hanzhi Yin, Yi Wang, Yifan He 외

The success of deep learning requires high-quality annotated and massive data. However, the size and the quality of a dataset are usually a trade-off in practice, as data collection and cleaning are expensive and time-co…

Metric LearningSpeaker RecognitionSpeaker Verification

Some Robustness Properties of Label Cleaning

2025-09-14 · Chen Cheng, John Duchi arxiv

We demonstrate that learning procedures that rely on aggregated labels, e.g., label information distilled from noisy responses, enjoy robustness properties impossible without data cleaning. This robustness appears in sev…

SimJEB: Simulated Jet Engine Bracket Dataset

2021-05-07 · Eamon Whalen, Azariah Beyene, Caitlin Mueller

This paper introduces the Simulated Jet Engine Bracket Dataset (SimJEB): a new, public collection of crowdsourced mechanical brackets and accompanying structural simulations. SimJEB is applicable to a wide range of geome…