paper-with-me

Papers

Information-Consistent Language Model Recommendations through Group Relative Policy Optimization

2025-12-14 · Sonal Prabhune, Balaji Padmanabhan, Kaushik Dutta arxiv

Large Language Models (LLMs) are increasingly deployed in business-critical domains such as finance, education, healthcare, and customer support, where users expect consistent and reliable recommendations. Yet LLMs often exhibit variability when prompts are phrased with minor differences, even when semantically equivalent. Such inconsistency undermines trust, complicates compliance, and disrupts user experience. While personalization is desirable in certain contexts, many enterprise scenarios, such as HR onboarding, customer support, or policy disclosure, require invariant information delivery regardless of phrasing or prior conversational history. Existing approaches, including retrieval-augmented generation (RAG) and temperature tuning, improve factuality or reduce stochasticity, but cannot guarantee stability across equivalent prompts. In this paper, we propose a reinforcement learning framework based on Group Relative Policy Optimization (GRPO) to directly optimize for consistency. Unlike prior applications of GRPO, which have been limited to reasoning and code generation, we adapt GRPO to enforce the stability of information content across groups of semantically equivalent prompts. We introduce entropy-based helpfulness and stability rewards, treating prompt variants as groups and resetting conversational context to isolate phrasing effects. Experiments on investment and job recommendation tasks show that our GRPO-fine-tuned model reduces variability compared to the baseline LLM model. To our knowledge, this is a novel application of GRPO for aligning LLMs toward information consistency, reframing variability not as an acceptable feature of generative diversity, but as a correctable flaw in enterprise deployments.

📄 PDF Abstract BibTeX arXiv:2512.12858

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningCode Generation

Similar Papers 제목 키워드 기반

Unleashing the Power of Large Language Models for Group POI Recommendations

2024-11-20 · Jing Long, Liang Qu, Guanhua Ye, Tong Chen 외

Group Point-of-Interest (POI) recommendations aim to predict the next POI that satisfies the diverse preferences of a group of users. This task is more challenging than traditional individual POI recommendations due to c…

Large Language ModelSelf-Supervised Learning

Consistent Explainers or Unreliable Narrators? Understanding LLM-generated Group Recommendations

2025-07-18 · Cedric Waterschoot, Nava Tintarev, Francesco Barile arxiv

Large Language Models (LLMs) are increasingly being implemented as joint decision-makers and explanation generators for Group Recommender Systems (GRS). In this paper, we evaluate these recommendations and explanations b…

Stereotype or Personalization? User Identity Biases Chatbot Recommendations

2024-10-08 · Anjali Kantharuban, Jeremiah Milbauer, Emma Strubell, Graham Neubig

We demonstrate that when people use large language models (LLMs) to generate recommendations, the LLMs produce responses that reflect both what the user wants and who the user is. While personalized recommendations are o…

Chatbot

Applying the Affective Aware Pseudo Association Method to Enhance the Top-N Recommendations Distribution to Users in Group Emotion Recommender Systems

2021-02-08 · John Kalung Leung, Igor Griva, William G. Kennedy

Recommender Systems are a subclass of information retrieval systems, or more succinctly, a class of information filtering systems that seeks to predict how close is the match of the user's preference to a recommended ite…

Decision MakingInformation RetrievalRecommendation SystemsRetrieval

"You Gotta be a Doctor, Lin": An Investigation of Name-Based Bias of Large Language Models in Employment Recommendations

2024-06-18 · Huy Nghiem, John Prindle, Jieyu Zhao, Hal Daumé III

Social science research has shown that candidates with names indicative of certain races or genders often face discrimination in employment practices. Similarly, Large Language Models (LLMs) have demonstrated racial and …