paper-with-me

홈 › Papers

EdgeJury: Cross-Reviewed Small-Model Ensembles for Truthful Question Answering on Serverless Edge Inference

2025-12-29 · Aayush Kumar arxiv

Hallucinations hinder reliable question answering, especially in resource-constrained deployments where frontier-scale models or retrieval pipelines may be impractical. We present EdgeJury, a lightweight ensemble framework that improves truthfulness and robustness using only small instruction-tuned language models (3B-8B) suitable for serverless edge inference. EdgeJury orchestrates four stages: (1) parallel role-specialized generation, (2) anonymized cross-review with structured critiques and rankings, (3) chairman synthesis that integrates the strongest content while addressing flagged issues, and (4) claim-level consistency labeling based on inter-model agreement. On TruthfulQA (MC1), EdgeJury achieves 76.2% accuracy (95% CI: 72.8-79.6%), a +21.4% relative improvement over a single 8B baseline (62.8%), and outperforms standard baselines including self-consistency and majority voting under transparent compute accounting (total tokens and platform cost reported). On a 200-question adversarial EdgeCases set, EdgeJury yields +48.2% relative gains (95% CI: 44.0-52.4%). Manual analysis on 100 incorrect answers shows an approximately 55% reduction in factual hallucination errors versus the single-model baseline. Deployed on Cloudflare Workers AI, EdgeJury achieves 8.4 s median end-to-end latency, demonstrating that coordinated small-model ensembles can improve truthfulness on misconception-heavy QA benchmarks without external retrieval or proprietary large-model APIs.

📄 PDF Abstract BibTeX arXiv:2601.00850

Code (0)

등록된 구현이 없습니다.

Tasks

Question Answering

Similar Papers 제목 키워드 기반

Improving Counterfactual Truthfulness for Molecular Property Prediction through Uncertainty Quantification

2025-04-03 · Jonas Teufel, Annika Leinweber, Pascal Friederich

Explainable AI (xAI) interventions aim to improve interpretability for complex black-box models, not only to improve user trust but also as a means to extract scientific insights from high-performing predictive systems. …

counterfactualMolecular Property PredictionProperty PredictionUncertainty Quantification

Truth Knows No Language: Evaluating Truthfulness Beyond English

2025-02-13 · Blanca Calvo Figueras, Eneko Sagarzazu, Julen Etxaniz, Jeremy Barnes 외

We introduce a professionally translated extension of the TruthfulQA benchmark designed to evaluate truthfulness in Basque, Catalan, Galician, and Spanish. Truthfulness evaluations of large language models (LLMs) have pr…

InformativenessMachine TranslationMultiple-choiceTranslation+1

Let LLMs Judge Each Other: Multi-Agent Peer-Reviewed Reasoning for Medical Question Answering

2026-06-13 · Zaifu Zhan, Shuang Zhou, Rui Zhang arxiv

Objective: To enhance the accuracy, interpretability, and robustness of large language models (LLMs) in medical question answering (MedQA). Method: We designed a multi-agent peer-reviewed reasoning method in which multip…

Question Answering

The unreasonable effectiveness of small neural ensembles in high-dimensional brain

2018-09-20 · A. N. Gorban, V. A. Makarov, I. Y. Tyukin

Despite the widely-spread consensus on the brain complexity, sprouts of the single neuron revolution emerged in neuroscience in the 1970s. They brought many unexpected discoveries, including grandmother or concept cells …

Vocal Bursts Intensity Prediction

Ensemble Debates with Local Large Language Models for AI Alignment

2025-08-27 · Ephraiem Sarabamoun arxiv

As large language models (LLMs) take on greater roles in high-stakes decisions, alignment with human values is essential. Reliance on proprietary APIs limits reproducibility and broad participation. We study whether loca…