paper-with-me

홈 › Papers

Ideology as a Problem: Lightweight Logit Steering for Annotator-Specific Alignment in Social Media Analysis

2025-12-08 · Wei Xia, Haowen Tang, Luozheng Li arxiv

LLMs internally organize political ideology along low-dimensional structures that are partially, but not fully aligned with human ideological space. This misalignment is systematic, model specific, and measurable. We introduce a lightweight linear probe that both quantifies the misalignment and minimally corrects the output layer. This paper introduces a simple and efficient method for aligning models with specific user opinions. Instead of retraining the model, we calculated a bias score from its internal features and directly adjusted the final output probabilities. This solution is practical and low-cost and preserves the original reasoning power of the model.

📄 PDF Abstract BibTeX arXiv:2601.04207

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

What Sounds ``Right'' to Me? Experiential Factors in the Perception of Political Ideology

2021-04-01 · EACL 2021 2 · Qinlan Shen, Carolyn Rose

In this paper, we challenge the assumption that political ideology is inherently built into text by presenting an investigation into the impact of experiential factors on annotator perceptions of political ideology. We c…

LogitsCoder: Towards Efficient Chain-of-Thought Path Search via Logits Preference Decoding for Code Generation

2026-02-15 · Jizheng Chen, Weiming Zhang, Xinyi Dai, Weiwen Liu 외 arxiv

Code generation remains a challenging task that requires precise and structured reasoning. Existing Test Time Scaling (TTS) methods, including structured tree search, have made progress in exploring reasoning paths but s…

Code Generation

Does Topic Sentiment Cause Perceived Ideology? Comparing Human and LLM Annotations in Political News Articles

2026-06-04 · Upasana Chatterjee arxiv

We ask whether topic sentiment has a causal effect on perceived political ideology, and whether the answer depends on who assigns the ideology label. Using articles from AllSides, paired with shared sentiment annotations…

Steering Language Models Before They Speak: Logit-Level Interventions

2026-01-16 · Hyeseon An, Shinwoo Park, Hyundong Jin, Yo-Sub Han arxiv

Controllable generation requires language models to realize output characteristics such as reading level, politeness, and toxicity. Existing steering methods are often indirect, require access to internal activations, or…

OGLS-SD: On-Policy Self-Distillation with Outcome-Guided Logit Steering for LLM Reasoning

2026-05-12 · Yuxiao Yang, Xiaoyun Wang, Weitong Zhang arxiv

We study on-policy self-distillation (OPSD), where a language model improves its reasoning ability by distilling privileged teacher distributions along its own on-policy trajectories. Despite its promise, OPSD can suffer…

Mathematical Reasoning