paper-with-me

홈 › Papers

Scaling Point-in-Time Language Models

2026-04-24 · Bryan Kelly, Semyon Malamud, Johannes Schwab, Teng Andrea Xu arxiv

Large language models trained on unrestricted internet corpora inevitably embed information from the future, introducing lookahead bias that compromises the validity of backtests and causal inference in finance and the social sciences. Point-in-time language models--trained exclusively on text available up to each calendar date--eliminate this leakage by construction, but existing efforts typically produce models that lag substantially behind their unconstrained counterparts. We show that this performance gap can be substantially narrowed through scale. Training decoder-only transformers with up to 4 billion parameters on 1 trillion chronologically filtered tokens from FineWeb, we construct a sequence of monthly model checkpoints spanning 2013-2024. Across a range of common-sense reasoning and language understanding benchmarks, our models approach the performance of leading open-weight models of comparable size (e.g., Gemma-3-4B and LLaMA-7B) trained on temporally unrestricted data, although a performance gap remains on several tasks. Instruction fine-tuning via LoRA further improves downstream usability. We release the complete pipeline--including dataset construction, training infrastructure, and evaluation code--to enable reproducible point-in-time language modeling and to support research applications that require strict temporal validity.

📄 PDF Abstract BibTeX arXiv:2607.11889

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Linguistic Generalizability of Test-Time Scaling in Mathematical Reasoning

2025-02-24 · Guijin Son, Jiwoo Hong, Hyunwoo Ko, James Thorne

Scaling pre-training compute has proven effective for achieving mulitlinguality, but does the same hold for test-time scaling? In this work, we introduce MCLM, a multilingual math benchmark featuring competition-level pr…

MathMathematical Reasoning

Item Response Scaling Laws: A Measurement Theory Approach for Efficient and Generalizable Neural Scaling Estimation

2026-05-29 · Sang Truong, Yuheng Tu, Rylan Schaeffer, Sanmi Koyejo arxiv

Scaling laws provide a fundamental framework for understanding the performance of Language Models (LMs), yet deriving them requires prohibitively expensive evaluations across thousands of checkpoints or millions of infer…

Microscaling Floating Point Formats for Large Language Models

2025-10-02 · Marco Cococcioni, Dario Pagani, Federico Rossi arxiv

The increasing computational and memory demands of large language models (LLMs) necessitate innovative approaches to optimize resource usage without compromising performance. This paper leverages microscaling floating-po…

CodeScaler: Scaling Code LLM Training and Test-Time Inference via Reward Models

2026-02-04 · Xiao Zhu, Xinyu Zhou, Boyu Zhu, Hanxu Hu 외 arxiv

Reinforcement Learning from Verifiable Rewards (RLVR) has driven recent progress in code large language models by leveraging execution-based feedback from unit tests, but its scalability is fundamentally constrained by t…

Reinforcement LearningCode Generation

Stabilizing Recurrent Dynamics for Test-Time Scalable Latent Reasoning in Looped Language Models

2026-05-26 · Xiao-Wen Yang, Ziyu Han, Xi-Hua Zhang, Wen-Da Wei 외 arxiv

Looped Language Models (LoopLMs) enable efficient latent reasoning through depth recurrence, yet exhibit unreliable test-time scaling behavior: performance often peaks at a certain iteration depth and then collapses with…

Mathematical Reasoning