paper-with-me

홈 › Papers

SAFE: A Sparse Autoencoder-Based Framework for Robust Query Enrichment and Hallucination Mitigation in LLMs

2025-03-04 · Samir Abdaljalil, Filippo Pallucchini, Andrea Seveso, Hasan Kurban, Fabio Mercorio, Erchin Serpedin

Despite the state-of-the-art performance of Large Language Models (LLMs), these models often suffer from hallucinations, which can undermine their performance in critical applications. In this work, we propose SAFE, a novel method for detecting and mitigating hallucinations by leveraging Sparse Autoencoders (SAEs). While hallucination detection techniques and SAEs have been explored independently, their synergistic application in a comprehensive system, particularly for hallucination-aware query enrichment, has not been fully investigated. To validate the effectiveness of SAFE, we evaluate it on two models with available SAEs across three diverse cross-domain datasets designed to assess hallucination problems. Empirical results demonstrate that SAFE consistently improves query generation accuracy and mitigates hallucinations across all datasets, achieving accuracy improvements of up to 29.45%.

📄 PDF Abstract BibTeX arXiv:2503.03032

Code (0)

등록된 구현이 없습니다.

Tasks

Hallucination

Similar Papers 제목 키워드 기반

Data Enrichment for Symbolic Regression Using Diffusion Models

2026-05-31 · Simon De Reuver, Tamas Kristof Toth, Teddy Lazebnik arxiv

Symbolic regression (SR) offers a route to scientific discovery by converting observations into interpretable governing equations. However, despite its promise, its reliability degrades sharply when spatiotemporal measur…

Steered Generation via Gradient-Based Optimization on Sparse Query Features

2026-05-21 · Sumanta Bhattacharyya, Pedram Rooshenas arxiv

Latent steering exploits internal representations of Large Language Models (LLMs) to guide generation, yet interventions on dense states can entangle distinct semantic features. In this paper, we investigate attention qu…

SAFER: Probing Safety in Reward Models with Sparse Autoencoder

2025-07-01 · Wei Shi, Ziyuan Xie, Sihang Li, Xiang Wang arxiv

Reinforcement learning from human feedback (RLHF) is a key paradigm for aligning large language models (LLMs) with human values, yet the reward models at its core remain largely opaque. In this work, we present Sparse Au…

Reinforcement Learning

Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs

2026-04-10 · Jinqi Luo, Jinyu Yang, Tal Neiman, Lei Fan 외 arxiv

Multimodal Large Language Models (MLLMs) have been shown to be vulnerable to malicious queries that can elicit unsafe responses. Recent work uses prompt engineering, response classification, or finetuning to improve MLLM…

Prompt Engineering

Automated Data Enrichment using Confidence-Aware Fine-Grained Debate among Open-Source LLMs for Mental Health and Online Safety

2025-12-06 · Junyu Mao, Anthony Hills, Talia Tseriotou, Maria Liakata 외 arxiv

Real-world indicators play an important role in many natural language processing (NLP) applications, such as life-event for mental health analysis and risky behaviour for online safety, yet labelling such information in …