paper-with-me

Papers

LLM Cognitive Judgements Differ From Human

2023-07-20 · Sotiris Lamprinidis

Large Language Models (LLMs) have lately been on the spotlight of researchers, businesses, and consumers alike. While the linguistic capabilities of such models have been studied extensively, there is growing interest in investigating them as cognitive subjects. In the present work I examine GPT-3 and ChatGPT capabilities on an limited-data inductive reasoning task from the cognitive science literature. The results suggest that these models' cognitive judgements are not human-like.

📄 PDF Abstract BibTeX arXiv:2307.11787

Code (1)

sotlampr/llm-cognitive-judgements 공식 구현

Methods 이 논문이 사용한 방법론

15 Ways to Contact How can i speak to someone at Delta Airlines 설명 없음
Attention 설명 없음
Adam 설명 없음
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Cosine Annealing Cosine Annealing is a type of learning rate schedule that has the effect of starting with a large learning rate that is relatively rapidly decreased to a minimum value before…

Similar Papers 제목 키워드 기반

Investigating Context Effects in Similarity Judgements in Large Language Models

2024-08-20 · Sagar Uprety, Amit Kumar Jaiswal, Haiming Liu, Dawei Song

Large Language Models (LLMs) have revolutionised the capability of AI models in comprehending and generating natural language text. They are increasingly being used to empower and deploy agents in real-world scenarios, w…

Are Language Models Sensitive to Morally Irrelevant Distractors?

2026-02-10 · Andrew Shaw, Christina Hahn, Catherine Rasgaitis, Yash Mishra 외 arxiv

With the rapid uptake of large language models (LLMs) across high-stakes settings, it is becoming increasingly important to ensure that LLMs behave in ways that align with human values. Existing moral benchmarks for this…

Moral Scenarios

Using profiles of cognitive capability to assess AI suitability for workplace tasks

2026-08-26 · Jonathan Prunty, Marko Tešić, Patrick Quinn, José Hernández-Orallo 외 arxiv

Organisations deploying AI face a scoping problem: which tasks can be automated, which should remain with humans, and which are best shared between the two. Aggregate benchmark scores provide little insight into where sy…

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation

2024-11-20 · David Otero, Javier Parapar, Álvaro Barreiro

Offline evaluation of search systems depends on test collections. These benchmarks provide the researchers with a corpus of documents, topics and relevance judgements indicating which documents are relevant for each topi…

Information RetrievalRetrieval

Quantum-like Structure in Multidimensional Relevance Judgements

2020-01-20 · Sagar Uprety, Prayag Tiwari, Shahram Dehdashti, Lauren Fell 외

A large number of studies in cognitive science have revealed that probabilistic outcomes of certain human decisions do not agree with the axioms of classical probability theory. The field of Quantum Cognition provides an…

Decision MakingDecision Making Under Uncertainty