paper-with-me

홈 › Papers

Benchmarking Large Language Models for Quebec Insurance: From Closed-Book to Retrieval-Augmented Generation

2026-03-08 · David Beauchemin, Richard Khoury arxiv

The digitization of insurance distribution in the Canadian province of Quebec, accelerated by legislative changes such as Bill 141, has created a significant "advice gap", leaving consumers to interpret complex financial contracts without professional guidance. While Large Language Models (LLMs) offer a scalable solution for automated advisory services, their deployment in high-stakes domains hinges on strict legal accuracy and trustworthiness. In this paper, we address this challenge by introducing AEPC-QA, a private gold-standard benchmark of 807 multiple-choice questions derived from official regulatory certification (paper) handbooks. We conduct a comprehensive evaluation of 51 LLMs across two paradigms: closed-book generation and retrieval-augmented generation (RAG) using a specialized corpus of Quebec insurance documents. Our results reveal three critical insights: 1) the supremacy of inference-time reasoning, where models leveraging chain-of-thought processing (e.g. o3-2025-04-16, o1-2024-12-17) significantly outperform standard instruction-tuned models; 2) RAG acts as a knowledge equalizer, boosting the accuracy of models with weak parametric knowledge by over 35 percentage points, yet paradoxically causing "context distraction" in others, leading to catastrophic performance regressions; and 3) a "specialization paradox", where massive generalist models consistently outperform smaller, domain-specific French fine-tuned ones. These findings suggest that while current architectures approach expert-level proficiency (~79%), the instability introduced by external context retrieval necessitates rigorous robustness calibration before autonomous deployment is viable.

📄 PDF Abstract BibTeX arXiv:2603.07825

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Quebec Automobile Insurance Question-Answering With Retrieval-Augmented Generation

2024-10-12 · David Beauchemin, Zachary Gagnon, Ricahrd Khoury

Large Language Models (LLMs) perform outstandingly in various downstream tasks, and the use of the Retrieval-Augmented Generation (RAG) architecture has been shown to improve performance for legal question answering (Nur…

Question AnsweringRAGRetrievalRetrieval-augmented Generation

RISC: Generating Realistic Synthetic Bilingual Insurance Contract

2023-04-09 · David Beauchemin, Richard Khoury

This paper presents RISC, an open-source Python package data generator (https://github.com/GRAAL-Research/risc). RISC generates look-alike automobile insurance contracts based on the Quebec regulatory insurance form in F…

Machine TranslationNERQuestion AnsweringText Simplification

InsQABench: Benchmarking Chinese Insurance Domain Question Answering with Large Language Models

2025-01-19 · Jing Ding, Kai Feng, Binbin Lin, Jiarui Cai 외

The application of large language models (LLMs) has achieved remarkable success in various fields, but their effectiveness in specialized domains like the Chinese insurance industry remains underexplored. The complexity …

BenchmarkingQuestion AnsweringRAG

Normalized insured losses caused by windstorms in Quebec and Ontario, Canada, in the period 2008-2021

2023-08-10 · Mohammad Hadavi, Lutong Sun, Djordje Romanic

Severe windstorms pose threats to people, human-made structures, and the environment. An investigation of insured losses caused by windstorms is a multipurpose study that serves to advance the resilience and sustainabili…

INS-MMBench: A Comprehensive Benchmark for Evaluating LVLMs' Performance in Insurance

2024-06-13 · Chenwei Lin, Hanjia Lyu, Xian Xu, Jiebo Luo

Large Vision-Language Models (LVLMs) have demonstrated outstanding performance in various general multimodal applications such as image recognition and visual reasoning, and have also shown promising potential in special…

Multiple-choiceVisual Reasoning