paper-with-me

Papers

DateLogicQA: Benchmarking Temporal Biases in Large Language Models

2024-12-17 · Gagan Bhatia, MingZe Tang, Cristina Mahanta, Madiha Kazi

This paper introduces DateLogicQA, a benchmark with 190 questions covering diverse date formats, temporal contexts, and reasoning types. We propose the Semantic Integrity Metric to assess tokenization quality and analyse two biases: Representation-Level Bias, affecting embeddings, and Logical-Level Bias, influencing reasoning outputs. Our findings provide a comprehensive evaluation of LLMs' capabilities and limitations in temporal reasoning, highlighting key challenges in handling temporal data accurately.

📄 PDF Abstract BibTeX arXiv:2412.13377

Code (1)

gagan3012/eais-temporal-bias 공식 구현

Tasks

Benchmarking

Similar Papers 제목 키워드 기반

Benchmarking Adversarial Robustness to Bias Elicitation in Large Language Models: Scalable Automated Assessment with LLM-as-a-Judge

2025-04-10 · Riccardo Cantini, Alessio Orsino, Massimo Ruggiero, Domenico Talia

Large Language Models (LLMs) have revolutionized artificial intelligence, driving advancements in machine translation, summarization, and conversational agents. However, their increasing integration into critical societa…

Adversarial RobustnessBenchmarkingFairnessMachine Translation

Social Bias Probing: Fairness Benchmarking for Language Models

2023-11-15 · Marta Marchiori Manerba, Karolina Stańczak, Riccardo Guidotti, Isabelle Augenstein

While the impact of social biases in language models has been recognized, prior methods for bias evaluation have been limited to binary association tests on small datasets, limiting our understanding of bias complexities…

BenchmarkingFairnessProbing Language Models

Social Bias in Large Language Models For Bangla: An Empirical Study on Gender and Religious Bias

2024-07-03 · Jayanta Sadhu, Maneesha Rani Saha, Rifat Shahriyar

The rapid growth of Large Language Models (LLMs) has put forward the study of biases as a crucial field. It is important to assess the influence of different types of biases embedded in LLMs to ensure fair use in sensiti…

BenchmarkingBias Detection

BOLD: Dataset and Metrics for Measuring Biases in Open-Ended Language Generation

2021-01-27 · Jwala Dhamala, Tony Sun, Varun Kumar, Satyapriya Krishna 외

Recent advances in deep learning techniques have enabled machines to generate cohesive open-ended text when prompted with a sequence of words as context. While these models now empower many downstream applications from c…

BenchmarkingText Generation

IndiBias: A Benchmark Dataset to Measure Social Biases in Language Models for Indian Context

2024-03-29 · Nihar Ranjan Sahoo, Pranamya Prashant Kulkarni, Narjis Asad, Arif Ahmad 외

The pervasive influence of social biases in language data has sparked the need for benchmark datasets that capture and evaluate these biases in Large Language Models (LLMs). Existing efforts predominantly focus on Englis…

BenchmarkingSentence