paper-with-me

홈 › Papers

Position: Measure Dataset Diversity, Don't Just Claim It

2024-07-11 · Dora Zhao, Jerone T. A. Andrews, Orestis Papakyriakopoulos, Alice Xiang

Machine learning (ML) datasets, often perceived as neutral, inherently encapsulate abstract and disputed social constructs. Dataset curators frequently employ value-laden terms such as diversity, bias, and quality to characterize datasets. Despite their prevalence, these terms lack clear definitions and validation. Our research explores the implications of this issue by analyzing "diversity" across 135 image and text datasets. Drawing from social sciences, we apply principles from measurement theory to identify considerations and offer recommendations for conceptualizing, operationalizing, and evaluating diversity in datasets. Our findings have broader implications for ML research, advocating for a more nuanced and precise approach to handling value-laden properties in dataset construction.

📄 PDF Abstract BibTeX arXiv:2407.08188

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityPosition

Similar Papers 제목 키워드 기반

Fact2Fiction: Targeted Poisoning Attack to Agentic Fact-checking System

2025-08-08 · Haorui He, Yupeng Li, Bin Benjamin Zhu, Dacheng Wen 외 arxiv

State-of-the-art (SOTA) fact-checking systems combat misinformation by employing autonomous LLM-based agents to decompose complex claims into smaller sub-claims, verify each sub-claim individually, and aggregate the part…

Do LLM Debates Repeat Arguments Differently Across Languages?

2026-07-26 · Huiqian Lai arxiv

LLM debate is usually evaluated by final answers, yet transcripts reveal whether later turns develop new arguments or return to earlier claims in new wording. We study this process with \textit{prior-argument similarity}…

Measuring Patent Claim Generation by Span Relevancy

2019-08-26 · Jieh-Sheng Lee, Jieh Hsiang

Our goal of patent claim generation is to realize "augmented inventing" for inventors by leveraging latest Deep Learning techniques. We envision the possibility of building an "auto-complete" function for inventors to co…

Language ModellingNatural Language InferenceText Generation

What is "Typological Diversity" in NLP?

2024-02-06 · Esther Ploeger, Wessel Poelman, Miryam de Lhoneux, Johannes Bjerva

The NLP research community has devoted increased attention to languages beyond English, resulting in considerable improvements for multilingual NLP. However, these improvements only apply to a small subset of the world's…

DiversityMultilingual NLP

A note on VIX for postprocessing quantitative strategies

2022-07-08 · Jun Lu, Minhui Wu

In this note, we introduce how to use Volatility Index (VIX) for postprocessing quantitative strategies so as to increase the Sharpe ratio and reduce trading risks. The signal from this procedure is an indicator of tradi…