paper-with-me

홈 › Papers

What Really Controls Temporal Reasoning in Large Language Models: Tokenisation or Representation of Time?

2026-03-19 · Gagan Bhatia, Ahmad Muhammad Isa, Maxime Peyrard, Wei Zhao arxiv

We present MultiTempBench, a multilingual temporal reasoning benchmark spanning three tasks, date arithmetic, time zone conversion, and temporal relation extraction across five languages (English, German, Chinese, Arabic, and Hausa) and multiple calendar conventions (Gregorian, Hijri, and Chinese Lunar). MultiTempBench contains $15,000$ examples built by translating $750$ curated English questions and expanding each into controlled date-format variants. We evaluate 20 LLMs and introduce the multilingual Date Fragmentation Ratio (mDFR), calibrated with human severity ratings, together with geometric-probing analyses of internal temporal representations. We find tokenisation quality of temporal artefacts is a resource-dependent bottleneck: in low-resource languages and rarer calendar formats, fragmentation disrupts Year/Month/Day separation and accuracy collapses, while high-resource settings are often robust to digit-level splitting. Beyond tokenisation, crossed mixed-effects regression shows that temporal linearity is the strongest predictor of temporal reasoning in high-resource languages, whereas fragmentation is the stronger predictor in low-resource languages. Code is available at: https://github.com/gagan3012/mtb

📄 PDF Abstract BibTeX arXiv:2603.19017

Code (0)

등록된 구현이 없습니다.

Tasks

Temporal Relation Extraction

Similar Papers 제목 키워드 기반

What is needed for simple spatial language capabilities in VQA?

2019-08-17 · Alexander Kuhnle, Ann Copestake

Visual question answering (VQA) comprises a variety of language capabilities. The diagnostic benchmark dataset CLEVR has fueled progress by helping to better assess and distinguish models in basic abilities like counting…

DiagnosticQuestion AnsweringSpatial ReasoningVisual Question Answering+1

What Really is Deep Learning Doing?

2017-11-06 · Chuyu Xiong

Deep learning has achieved a great success in many areas, from computer vision to natural language processing, to game playing, and much more. Yet, what deep learning is really doing is still an open question. There are …

Deep LearningOpen-Ended Question Answering

Vinoground: Scrutinizing LMMs over Dense Temporal Reasoning with Short Videos

2024-10-03 · Jianrui Zhang, Mu Cai, Yong Jae Lee

There has been growing sentiment recently that modern large multimodal models (LMMs) have addressed most of the key challenges related to short video comprehension. As a result, both academia and industry are gradually s…

counterfactual

What's the Meaning of Superhuman Performance in Today's NLU?

2023-05-15 · Simone Tedeschi, Johan Bos, Thierry Declerck, Jan Hajic 외

In the last five years, there has been a significant focus in Natural Language Processing (NLP) on developing larger Pretrained Language Models (PLMs) and introducing benchmarks such as SuperGLUE and SQuAD to measure the…

PositionReading Comprehension

What Can We Really Learn from Post-editing?

2016-10-01 · AMTA 2016 10 · Marcis Pinnis, Rihards Kalnins, Raivis Skadins, Inguna Skadina