paper-with-me

홈 › Papers

Chronologically Consistent Large Language Models

2025-02-28 · Songrun He, Linying Lv, Asaf Manela, Jimmy Wu

Large language models are increasingly used in social sciences, but their training data can introduce lookahead bias and training leakage. A good chronologically consistent language model requires efficient use of training data to maintain accuracy despite time-restricted data. Here, we overcome this challenge by training a suite of chronologically consistent large language models, ChronoBERT and ChronoGPT, which incorporate only the text data that would have been available at each point in time. Despite this strict temporal constraint, our models achieve strong performance on natural language processing benchmarks, outperforming or matching widely used models (e.g., BERT), and remain competitive with larger open-weight models. Lookahead bias is model and application-specific because even if a chronologically consistent language model has poorer language comprehension, a regression or prediction model applied on top of the language model can compensate. In an asset pricing application predicting next-day stock returns from financial news, we find that ChronoBERT's real-time outputs achieve a Sharpe ratio comparable to state-of-the-art models, indicating that lookahead bias is modest. Our results demonstrate a scalable, practical framework to mitigate training leakage, ensuring more credible backtests and predictions across finance and other social science domains.

📄 PDF Abstract BibTeX arXiv:2502.21206

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modelling

Methods 이 논문이 사용한 방법론

Lookahead 설명 없음

Similar Papers 제목 키워드 기반

Instruction Tuning Chronologically Consistent Language Models

2025-10-13 · Songrun He, Linying Lv, Asaf Manela, Jimmy Wu arxiv

We introduce a family of chronologically consistent, instruction-tuned large language models to eliminate lookahead bias. Each model is trained only on data available before a clearly defined knowledge-cutoff date, ensur…

Chronologically Accurate Retrieval for Temporal Grounding of Motion-Language Models

2024-07-22 · Kent Fujiwara, Mikihiro Tanaka, Qing Yu

With the release of large-scale motion datasets with textual annotations, the task of establishing a robust latent space for language and 3D human motion has recently witnessed a surge of interest. Methods have been prop…

Motion GenerationRetrieval

Temporal Analysis of Language through Neural Language Models

2014-05-14 · WS 2014 6 · Yoon Kim, Yi-I Chiu, Kentaro Hanaki, Darshan Hegde 외

We provide a method for automatically detecting change in language across time through a chronologically trained neural language model. We train the model on the Google Books Ngram corpus to obtain word vector representa…

Language ModelingLanguage Modelling

A Web Interface for Diachronic Semantic Search in Spanish

2017-04-01 · EACL 2017 4 · Pablo Gamallo, Iv{\'a}n Rodr{\'\i}guez-Torres, Marcos Garcia

This article describes a semantic system which is based on distributional models obtained from a chronologically structured language resource, namely Google Books Syntactic Ngrams.The models were created using dependency…

Projecting named entity recognizers without annotated or parallel corpora

2019-09-01 · WS (NoDaLiDa) 2019 9 · Jue Hou, Maximilian Koppatz, José María Hoya Quecedo, Roman Yangarber

Named entity recognition (NER) is a well-researched task in the field of NLP, which typically requires large annotated corpora for training usable models. This is a problem for languages which lack large annotated corpor…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER