paper-with-me

홈 › Papers

VERITAS: A Unified Approach to Reliability Evaluation

2024-11-05 · Rajkumar Ramamurthy, Meghana Arakkal Rajeev, Oliver Molenschot, James Zou, Nazneen Rajani

Large language models (LLMs) often fail to synthesize information from their context to generate an accurate response. This renders them unreliable in knowledge intensive settings where reliability of the output is key. A critical component for reliable LLMs is the integration of a robust fact-checking system that can detect hallucinations across various formats. While several open-access fact-checking models are available, their functionality is often limited to specific tasks, such as grounded question-answering or entailment verification, and they perform less effectively in conversational settings. On the other hand, closed-access models like GPT-4 and Claude offer greater flexibility across different contexts, including grounded dialogue verification, but are hindered by high costs and latency. In this work, we introduce VERITAS, a family of hallucination detection models designed to operate flexibly across diverse contexts while minimizing latency and costs. VERITAS achieves state-of-the-art results considering average performance on all major hallucination detection benchmarks, with $10\%$ increase in average performance when compared to similar-sized models and get close to the performance of GPT4 turbo with LLM-as-a-judge setting.

📄 PDF Abstract BibTeX arXiv:2411.03300

Code (0)

등록된 구현이 없습니다.

Tasks

Fact CheckingHallucinationQuestion Answering

Methods 이 논문이 사용한 방법론

Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Position-Wise Feed-Forward Layer 설명 없음
Adam 설명 없음
Attention 설명 없음
Multi-Head Attention 설명 없음
Residual Connection 설명 없음
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…

Similar Papers 제목 키워드 기반

VeriTaS: The First Dynamic Benchmark for Multimodal Automated Fact-Checking

2026-01-13 · Mark Rothermel, Marcus Kornmann, Marcus Rohrbach, Anna Rohrbach arxiv

The growing scale of online misinformation urgently demands Automated Fact-Checking (AFC). Existing benchmarks for evaluating AFC systems, however, are largely limited in terms of task scope, modalities, domain, language…

SVeritas: Benchmark for Robust Speaker Verification under Diverse Conditions

2025-09-21 · Massa Baali, Sarthak Bisht, Francisco Teixeira, Kateryna Shapovalenko 외 arxiv

Speaker verification (SV) models are increasingly integrated into security, personalization, and access control systems, yet their robustness to many real-world challenges remains inadequately benchmarked. These include …

Speaker Verification

Versatile Verification of Tree Ensembles

2020-10-26 · Laurens Devos, Wannes Meert, Jesse Davis

Machine learned models often must abide by certain requirements (e.g., fairness or legal). This has spurred interested in developing approaches that can provably verify whether a model satisfies certain properties. This …

Fairness

VERITAS: Towards a General-Purpose Replication Tool for Scientific Research

2026-07-03 · Haokun Liu, Filbert Aurelian Tjiaranata, Chenhao Tan arxiv

AI tools are accelerating scientific publication while the systems that review it struggle to keep up, and independent verification of published research has become both harder and more important. As manual replication i…

VERITAS: Verifying the Performance of AI-native Transceiver Actions in Base-Stations

2025-01-01 · Nasim Soltani, Michael Loehning, Kaushik Chowdhury

Artificial Intelligence (AI)-native receivers prove significant performance improvement in high noise regimes and can potentially reduce communication overhead compared to the traditional receiver. However, their perform…