paper-with-me

홈 › Papers

TRAM: Benchmarking Temporal Reasoning for Large Language Models

2023-10-02 · Yuqing Wang, Yun Zhao

Reasoning about time is essential for understanding the nuances of events described in natural language. Previous research on this topic has been limited in scope, characterized by a lack of standardized benchmarks that would allow for consistent evaluations across different studies. In this paper, we introduce TRAM, a temporal reasoning benchmark composed of ten datasets, encompassing various temporal aspects of events such as order, arithmetic, frequency, and duration, designed to facilitate a comprehensive evaluation of the TeR capabilities of large language models (LLMs). We evaluate popular LLMs like GPT-4 and Llama2 in zero-shot and few-shot scenarios, and establish baselines with BERT-based and domain-specific models. Our findings indicate that the best-performing model lags significantly behind human performance. It is our aspiration that TRAM will spur further progress in enhancing the TeR capabilities of LLMs.

📄 PDF Abstract BibTeX arXiv:2310.00835

Code (1)

eternityyw/tram-benchmark 공식 구현

Tasks

BenchmarkingFew-Shot Learning

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Adam 설명 없음
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…

Similar Papers 제목 키워드 기반

DateLogicQA: Benchmarking Temporal Biases in Large Language Models

2024-12-17 · Gagan Bhatia, MingZe Tang, Cristina Mahanta, Madiha Kazi

This paper introduces DateLogicQA, a benchmark with 190 questions covering diverse date formats, temporal contexts, and reasoning types. We propose the Semantic Integrity Metric to assess tokenization quality and analyse…

Benchmarking

Towards Benchmarking and Improving the Temporal Reasoning Capability of Large Language Models

2023-06-15 · Qingyu Tan, Hwee Tou Ng, Lidong Bing

Reasoning about time is of fundamental importance. Many facts are time-dependent. For example, athletes change teams from time to time, and different government officials are elected periodically. Previous time-dependent…

BenchmarkingQuestion Answering

Time Series Generation with Masked Autoencoder

2022-01-14 · Mengyue Zha, SiuTim Wong, Mengqi Liu, Tong Zhang 외

This paper shows that masked autoencoder with extrapolator (ExtraMAE) is a scalable self-supervised model for time series generation. ExtraMAE randomly masks some patches of the original time series and learns temporal d…

Data AugmentationDecoderImputationManagement+5

Benchmarking Temporal Reasoning and Alignment Across Chinese Dynasties

2025-02-24 · Zhenglin Wang, Jialong Wu, Pengfei Li, Yong Jiang 외

Temporal reasoning is fundamental to human cognition and is crucial for various real-world applications. While recent advances in Large Language Models have demonstrated promising capabilities in temporal reasoning, exis…

Benchmarking

eTraM: Event-based Traffic Monitoring Dataset

2024-03-29 · CVPR 2024 1 · Aayush Atul Verma, Bharatesh Chakravarthi, Arpitsinh Vaghela, Hua Wei 외

Event cameras, with their high temporal and dynamic range and minimal memory usage, have found applications in various fields. However, their potential in static traffic monitoring remains largely unexplored. To facilita…