Probabilistic Aggregation and Targeted Embedding Optimization for Collective Moral Reasoning in Large Language Models
Large Language Models (LLMs) have shown impressive moral reasoning abilities. Yet they often diverge when confronted with complex, multi-factor moral dilemmas. To address these discrepancies, we propose a framework that synthesizes multiple LLMs' moral judgments into a collectively formulated moral judgment, realigning models that deviate significantly from this consensus. Our aggregation mechanism fuses continuous moral acceptability scores (beyond binary labels) into a collective probability, weighting contributions by model reliability. For misaligned models, a targeted embedding-optimization procedure fine-tunes token embeddings for moral philosophical theories, minimizing JS divergence to the consensus while preserving semantic integrity. Experiments on a large-scale social moral dilemma dataset show our approach builds robust consensus and improves individual model fidelity. These findings highlight the value of data-driven moral alignment across multiple models and its potential for safer, more consistent AI systems.
Code (1)
Similar Papers 제목 키워드 기반
DeepVoting: Learning Voting Rules with Tailored Embeddings
Aggregating the preferences of multiple agents into a collective decision is a common step in many important problems across areas of computer science including information retrieval, reinforcement learning, and recommen…
Information RetrievalRecommendation SystemsA model aggregation approach for high-dimensional large-scale optimization
Bayesian optimization (BO) has been widely used in machine learning and simulation optimization. With the increase in computational resources and storage capacities in these fields, high-dimensional and large-scale probl…
Bayesian OptimizationFace DetectionVocal Bursts Intensity PredictionAggregating Credences into Beliefs: Agenda Conditions for Impossibility Results
Binarizing belief aggregation addresses how to rationally aggregate individual probabilistic beliefs into collective binary beliefs. Similar to the development of judgment aggregation theory, formulating axiomatic requir…
BinarizationNegationBounded arbitrage and nearly rational behavior
We establish the equivalence between a principle of almost absence of arbitrage opportunities and nearly rational decision-making. The implications of such principle are considered in the context of the aggregation of pr…
Decision MakingBeyond Node Embedding: A Direct Unsupervised Edge Representation Framework for Homogeneous Networks
Network representation learning has traditionally been used to find lower dimensional vector representations of the nodes in a network. However, there are very important edge driven mining tasks of interest to the classi…
Link PredictionNetwork EmbeddingRepresentation Learning