paper-with-me

Papers

Harnessing Temporal Databases for Systematic Evaluation of Factual Time-Sensitive Question-Answering in Large Language Models

2025-08-04 · Soyeon Kim, Jindong Wang, Xing Xie, Steven Euijong Whang arxiv

Facts change over time, making it essential for Large Language Models (LLMs) to handle time-sensitive factual knowledge accurately and reliably. Although factual Time-Sensitive Question-Answering (TSQA) tasks have been widely developed, existing benchmarks often face manual bottlenecks that limit scalable and comprehensive TSQA evaluation. To address this issue, we propose TDBench, a new benchmark that systematically constructs TSQA pairs by harnessing temporal databases and database techniques, such as temporal functional dependencies, temporal SQL, and temporal joins. We also introduce a new evaluation metric called time accuracy, which assesses the validity of time references in model explanations alongside traditional answer accuracy for a more fine-grained TSQA evaluation. Extensive experiments on contemporary LLMs show how TDBench enables scalable and comprehensive TSQA evaluation while reducing the reliance on human labor, complementing current TSQA evaluation approaches that largely center on Wikipedia/Wikidata by enabling LLM evaluation on application-specific data.

📄 PDF Abstract BibTeX arXiv:2508.02045

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Systematic Evaluation of the Quality of Synthetic Clinical Notes Rephrased by LLMs at Million-Note Scale

2026-05-18 · Jinghui Liu, Sarvesh Soni, Anthony Nguyen arxiv

Large language models (LLMs) can generate or synthesize clinical text for a wide range of applications, from improving clinical documentation to augmenting clinical text analytics. Yet evaluations typically focus on a na…

When Benchmarks Age: Temporal Misalignment through Large Language Model Factuality Evaluation

2025-10-08 · Xunyi Jiang, Dingyi Chang, Julian McAuley, Xin Xu arxiv

The rapid evolution of large language models (LLMs) and the real world has outpaced the static nature of widely used evaluation benchmarks, raising concerns about their reliability for evaluating LLM factuality. While su…

DREAM: Deep Research Evaluation with Agentic Metrics

2026-02-21 · Elad Ben Avraham, Changhao Li, Ron Dorfman, Roy Ganz 외 arxiv

Deep Research Agents generate analyst-grade reports, yet evaluating them remains challenging due to the absence of a single ground truth and the multidimensional nature of research quality. Recent benchmarks propose dist…

Answer-Set Programs for Repair Updates and Counterfactual Interventions

2022-09-25 · Leopoldo Bertossi

We briefly describe -- mainly through very simple examples -- different kinds of answer-set programs with annotations that have been proposed for specifying: database repairs and consistent query answering; secrecy view …

counterfactual

Data mining of public genomic repositories: harnessing off-target reads to expand microbial pathogen genomic resources

2025-05-15 · Damien Richard, Nils Poulicard

As sequencing technologies become more affordable and genomic databases expand continuously, the reuse of publicly available sequencing data emerges as a powerful strategy for studying microbial pathogens. Indeed, raw se…