paper-with-me

홈 › Papers

Temporal Task Diversity: Inductive Biases Under Non-Stationarity in Synthetic Sequence Modelling

2026-05-18 · Afiq Abdillah Effiezal Aswadi, Oliver Britton, Ross Baker, Matthew Farrugia-Roberts arxiv

Modern deep learning science often assumes that neural networks learn from a fixed data distribution. However, many practically important learning problems involve data distributions that change throughout training. How does such non-stationarity impact the inductive biases of deep learning towards models with different structural, generalisation, and safety properties? A fruitful testbed for studying inductive bias is in-context linear regression sequence modelling, where small transformers display strikingly different generalisation patterns depending on the diversity of the (fixed) training task distribution. In this paper, we explore the effect of diversifying the task distribution across training time, finding that such temporal diversity leads to an increased bias towards generalisation over memorisation.

📄 PDF Abstract BibTeX arXiv:2605.18281

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

You Don't Need Strong Assumptions: Visual Representation Learning via Temporal Differences

2026-06-14 · Ninad Daithankar, Alexi Gladstone, Yann LeCun, Heng Ji arxiv

Progress in AI has largely been driven by methods that assume less. As compute and data increase, approaches with weaker inductive biases generally outperform those with stronger assumptions. This is particularly charact…

Self-Supervised LearningRepresentation Learning

Drift Happens: An Empirical Study of Neural Architecture Robustness to Temporal Distribution Shift

2026-07-07 · Robin Holzinger, Riccardo Colletti arxiv

Real-world data distributions evolve over time, inducing temporal distribution shift that can substantially degrade the reliability of deployed machine learning systems. However, the extent to which architectural choices…

Multi-Label Text ClassificationImage Classification

Integrating Inductive Biases in Transformers via Distillation for Financial Time Series Forecasting

2026-03-17 · Yu-Chen Den, Kuan-Yu Chen, Kendro Vincent, Darby Tien-Hao Chang arxiv

Transformer-based models have been widely adopted for time-series forecasting due to their high representational capacity and architectural flexibility. However, many Transformer variants implicitly assume stationarity a…

Time Series ForecastingKnowledge Distillation

Video Understanding by Design: How Datasets Shape Video Models

2025-09-11 · Lei Wang, Syuan-Hao Li, Piotr Koniusz, Yongsheng Gao arxiv

Research in video understanding has advanced rapidly, driven by increasingly diverse datasets and more powerful model architectures. While existing surveys typically organize progress by tasks, benchmarks, or model famil…

Biases for Emergent Communication in Multi-agent Reinforcement Learning

2019-12-11 · NeurIPS 2019 12 · Tom Eccles, Yoram Bachrach, Guy Lever, Angeliki Lazaridou 외

We study the problem of emergent communication, in which language arises because speakers and listeners must communicate information in order to solve tasks. In temporally extended reinforcement learning domains, it has …

Multi-agent Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)