paper-with-me

Papers

Aligning Large Language Model Behavior with Human Citation Preferences

2026-02-05 · Kenichiro Ando, Tatsuya Harada arxiv

Most services built on powerful large-scale language models (LLMs) add citations to their output to enhance credibility. Recent research has paid increasing attention to the question of what reference documents to link to outputs. However, how LLMs recognize cite-worthiness and how this process should be controlled remains underexplored. In this study, we focus on what kinds of content LLMs currently tend to cite and how well that behavior aligns with human preferences. We construct a dataset to characterize the relationship between human citation preferences and LLM behavior. Web-derived texts are categorized into eight citation-motivation types, and pairwise citation preferences are exhaustively evaluated across all type combinations to capture fine-grained contrasts. Our results show that humans most frequently seek citations for medical text, and stronger models display a similar tendency. We also find that current models are as much as $27\%$ more likely than humans to add citations to text that is explicitly marked as needing citations on sources such as Wikipedia, and this overemphasis reduces alignment accuracy. Conversely, models systematically underselect numeric sentences (by $-22.6\%$ relative to humans) and sentences containing personal names (by $-20.1\%$), categories for which humans typically demand citations. Furthermore, experiments with Direct Preference Optimization demonstrate that model behavior can be calibrated to better match human citation preferences. We expect this study to provide a foundation for more fine-grained investigations into LLM citation preferences.

📄 PDF Abstract BibTeX arXiv:2602.05205

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Eliciting Human Preferences with Language Models

2023-10-17 · Belinda Z. Li, Alex Tamkin, Noah Goodman, Jacob Andreas

Language models (LMs) can be directed to perform target tasks by using labeled examples or natural language prompts. But selecting examples or writing prompts for can be challenging--especially in tasks that involve unus…

Incentivizing Truthful Language Models via Peer Elicitation Games

2025-05-19 · Baiting Chen, Tong Zhu, Jiale Han, Lexin Li 외

Large Language Models (LLMs) have demonstrated strong generative capabilities but remain prone to inconsistencies and hallucinations. We introduce Peer Elicitation Games (PEG), a training-free, game-theoretic framework f…

ReviewGuard: Aligning LLM-Assisted Peer Review with Long-Term Scientific Impact

2026-05-29 · Abdur Rasool, Xiaohui Huang, Yanqing Hu, Linyi Yang arxiv

Peer review is central to scientific quality control, yet it can undervalue papers that later achieve substantial citation impact. While frontier large language models have shown promise in automating aspects of peer rev…

Reinforcement Learning

Aligning Agents like Large Language Models

2024-06-06 · Adam Jelley, Yuhan Cao, Dave Bignell, Sam Devlin 외

Training agents to behave as desired in complex 3D environments from high-dimensional sensory information is challenging. Imitation learning from diverse human behavior provides a scalable approach for training an agent …

Imitation Learning

Aligned Textual Scoring Rules

2025-07-08 · Yuxuan Lu, Yifan Wu, Jason Hartline, Michael J. Curry

Scoring rules elicit probabilistic predictions from a strategic agent by scoring the prediction against a ground truth state. A scoring rule is proper if, from the agent's perspective, reporting the true belief maximizes…

scoring rule