paper-with-me

Papers

Are LLMs Prescient? A Continuous Evaluation using Daily News as the Oracle

2024-11-13 · Hui Dai, Ryan Teehan, Mengye Ren

Many existing evaluation benchmarks for Large Language Models (LLMs) quickly become outdated due to the emergence of new models and training data. These benchmarks also fall short in assessing how LLM performance changes over time, as they consist of static questions without a temporal dimension. To address these limitations, we propose using future event prediction as a continuous evaluation method to assess LLMs' temporal generalization and forecasting abilities. Our benchmark, Daily Oracle, automatically generates question-answer (QA) pairs from daily news, challenging LLMs to predict "future" event outcomes. Our findings reveal that as pre-training data becomes outdated, LLM performance degrades over time. While Retrieval Augmented Generation (RAG) has the potential to enhance prediction accuracy, the performance degradation pattern persists, highlighting the need for continuous model updates.

📄 PDF Abstract BibTeX arXiv:2411.08324

Code (0)

등록된 구현이 없습니다.

Tasks

RAGRetrievalRetrieval-augmented Generation

Similar Papers 제목 키워드 기반

ORACLE: Time-Dependent Recursive Summary Graphs for Foresight on News Data Using LLMs

2025-12-17 · Lev Kharlashkin, Eiaki Morooka, Yehor Tereshchenko, Mika Hämäläinen arxiv

ORACLE turns daily news into week-over-week, decision-ready insights for one of the Finnish University of Applied Sciences. The platform crawls and versions news, applies University-specific relevance filtering, embeds c…

Acquiring Predicate Paraphrases from News Tweets

2017-08-01 · SEMEVAL 2017 8 · Vered Shwartz, Gabriel Stanovsky, Ido Dagan

We present a simple method for ever-growing extraction of predicate paraphrases from news headlines in Twitter. Analysis of the output of ten weeks of collection shows that the accuracy of paraphrases with different supp…

Natural Language InferenceQuestion Answering

CN-Buzz2Portfolio: A Chinese-Market Dataset and Benchmark for LLM-Based Macro and Sector Asset Allocation from Daily Trending Financial News

2026-03-18 · Liyuan Chen, Shilong Li, Jiangpeng Yan, Shuoling Liu 외 arxiv

Large Language Models (LLMs) are rapidly transitioning from static Natural Language Processing (NLP) tasks including sentiment analysis and event extraction to acting as dynamic decision-making agents in complex financia…

Sentiment AnalysisEvent Extraction

The American Local News Corpus

2014-05-01 · LREC 2014 5 · Ann Irvine, Joshua Langfus, Chris Callison-Burch

We present the American Local News Corpus (ALNC), containing over 4 billion words of text from 2,652 online newspapers in the United States. Each article in the corpus is associated with a timestamp, state, and city. All…

FFN: a Fine-grained Chinese-English Financial Domain Parallel Corpus

2024-06-27 · Yuxin Fu, Shijing Si, Leyi Mai, Xi-ang Li

Large Language Models (LLMs) have stunningly advanced the field of machine translation, though their effectiveness within the financial domain remains largely underexplored. To probe this issue, we constructed a fine-gra…

ArticlesMachine TranslationTranslation