paper-with-me

Papers

IndoPref: A Multi-Domain Pairwise Preference Dataset for Indonesian

2025-07-29 · Vanessa Rebecca Wiyono, David Anugraha, Ayu Purwarianti, Genta Indra Winata arxiv

Over 200 million people speak Indonesian, yet the language remains significantly underrepresented in preference-based research for large language models (LLMs). Most existing multilingual datasets are derived from English translations, often resulting in content that lacks cultural and linguistic authenticity. To address this gap, we introduce IndoPref, the first fully human-authored and multi-domain Indonesian preference dataset designed to evaluate the naturalness and quality of LLM-generated text. The dataset contains 522 prompts and yields 4,099 human-annotated pairwise preferences from comparisons across five instruction-tuned LLMs. All annotations are natively written in Indonesian with strong inter-annotator agreement, measured by Krippendorff's alpha. Our benchmark spans 10 diverse categories, enabling practitioners to identify LLMs' fine-grained strengths and weaknesses.

📄 PDF Abstract BibTeX arXiv:2507.22159

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Preferences Order, Ratings Anchor: From Fused Expert Aesthetic Ground Truth to Self-Distillation

2026-05-19 · Yuanpei Zhao, Jie Lin, Chao Zhang, Yilin Wang 외 arxiv

Pairwise preferences and pointwise ratings are the two dominant annotation protocols in image aesthetic assessment (IAA), yet existing benchmarks adopt only one, leaving their complementarity unmeasured under controlled …

Estimating Summary Quality with Pairwise Preferences

2018-06-01 · NAACL 2018 6 · Markus Zopf

Automatic evaluation systems in the field of automatic summarization have been relying on the availability of gold standard summaries for over ten years. Gold standard summaries are expensive to obtain and often require …

Text Summarization

ToolRM: Towards Agentic Tool-Use Reward Modeling

2025-10-30 · Renhao Li, Jianhong Tu, Yang Su, Yantao Liu 외 arxiv

Reward models (RMs) play a critical role in aligning large language models (LLMs) with human preferences. Yet in the domain of tool learning, the lack of RMs specifically designed for function-calling tasks has limited p…

JudgmentBench: Comparing Rubric and Preference Evaluation for Quality Assessment

2026-05-24 · Russell Yang, Ruishi Chen, Pierce Kelaita, Riya Ranjan 외 arxiv

Two methodologies dominate current practices of benchmarking: rubric-based scoring evaluates items against predefined criteria, whereas comparative judgment elicits pairwise preferences between outputs. Although both met…

Computing Voting Rules with Improvement Feedback

2025-02-18 · Evi Micha, Vasilis Varsamis

Aggregating preferences under incomplete or constrained feedback is a fundamental problem in social choice and related domains. While prior work has established strong impossibility results for pairwise comparisons, this…