paper-with-me

Papers

SubData: Bridging Heterogeneous Datasets to Enable Theory-Driven Evaluation of Political and Demographic Perspectives in LLMs

2024-12-21 · Leon Fröhling, Pietro Bernardelle, Gianluca Demartini

As increasingly capable large language models (LLMs) emerge, researchers have begun exploring their potential for subjective tasks. While recent work demonstrates that LLMs can be aligned with diverse human perspectives, evaluating this alignment on actual downstream tasks (e.g., hate speech detection) remains challenging due to the use of inconsistent datasets across studies. To address this issue, in this resource paper we propose a two-step framework: we (1) introduce SubData, an open-source Python library designed for standardizing heterogeneous datasets to evaluate LLM perspective alignment; and (2) present a theory-driven approach leveraging this library to test how differently-aligned LLMs (e.g., aligned with different political viewpoints) classify content targeting specific demographics. SubData's flexible mapping and taxonomy enable customization for diverse research needs, distinguishing it from existing resources. We invite contributions to add datasets to our initially proposed resource and thereby help expand SubData into a multi-construct benchmark suite for evaluating LLM perspective alignment on NLP tasks.

📄 PDF Abstract BibTeX arXiv:2412.16783

Code (0)

등록된 구현이 없습니다.

Tasks

Hate Speech Detection

Methods 이 논문이 사용한 방법론

Library 설명 없음
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Nearly Optimal Subdata Selection

2026-04-27 · Min Yang, Wei Zheng, John Stufken, Ming-Chung Chang 외 arxiv

When, in terms of the number of data points, the size of a dataset exceeds available computing resources, or when labeling is expensive, an attractive solution consists of selecting only some of the data points (subdata)…

A model-free subdata selection method for classification

2024-04-29 · Rakhi Singh

Subdata selection is a study of methods that select a small representative sample of the big data, the analysis of which is fast and statistically efficient. The existing subdata selection methods assume that the big dat…

Classification

Model-free Subsampling Method Based on Uniform Designs

2022-09-08 · Mei Zhang, Yongdao Zhou, Zheng Zhou, Aijun Zhang

Subsampling or subdata selection is a useful approach in large-scale statistical learning. Most existing studies focus on model-based subsampling methods which significantly depend on the model assumption. In this paper,…

Hybrid Unet-Transformer Model for Generating Stress and Strain Fields from Composite Geometrics

2026-06-30 · Shrey Patel arxiv

Accurate prediction of stress and strain fields in hierarchical composite microstructures is critical for physics-informed material design, yet conventional finite element method (FEM) simulations are computationally pro…

Powering In-Database Dynamic Model Slicing for Structured Data Analytics

2024-05-01 · Lingze Zeng, Naili Xing, Shaofeng Cai, Gang Chen 외

Relational database management systems (RDBMS) are widely used for the storage of structured data. To derive insights beyond statistical aggregation, we typically have to extract specific subdatasets from the database us…

Mixture-of-Experts