paper-with-me

Papers

Does GPT Really Get It? A Hierarchical Scale to Quantify Human vs AI's Understanding of Algorithms

2024-06-20 · Mirabel Reid, Santosh S. Vempala

As Large Language Models (LLMs) perform (and sometimes excel at) more and more complex cognitive tasks, a natural question is whether AI really understands. The study of understanding in LLMs is in its infancy, and the community has yet to incorporate well-trodden research in philosophy, psychology, and education. We initiate this, specifically focusing on understanding algorithms, and propose a hierarchy of levels of understanding. We use the hierarchy to design and conduct a study with human subjects (undergraduate and graduate students) as well as large language models (generations of GPT), revealing interesting similarities and differences. We expect that our rigorous criteria will be useful to keep track of AI's progress in such cognitive domains.

📄 PDF Abstract BibTeX arXiv:2406.14722

Code (0)

등록된 구현이 없습니다.

Tasks

Philosophy

Similar Papers 제목 키워드 기반

Sequence Classification with Human Attention

2018-10-01 · CONLL 2018 10 · Maria Barrett, Joachim Bingel, Nora Hollenstein, Marek Rei 외

Learning attention functions requires large volumes of data, but many NLP tasks simulate human behavior, and in this paper, we show that human attention really does provide a good inductive bias on many attention functio…

Abusive LanguageClassificationGeneral ClassificationGrammatical Error Detection+2

On the Ethics of Building AI in a Responsible Manner

2020-03-30 · Shai Shalev-Shwartz, Shaked Shammah, Amnon Shashua

The AI-alignment problem arises when there is a discrepancy between the goals that a human designer specifies to an AI learner and a potential catastrophic outcome that does not reflect what the human designer really wan…

BIG-bench Machine LearningEthics

MT Quality Estimation for Computer-assisted Translation: Does it Really Help?

2015-07-01 · IJCNLP 2015 7 · Marco Turchi, Matteo Negri, Marcello Federico
Machine TranslationTranslation

When does word order matter and when doesn't it?

2024-02-29 · Xuanda Chen, Timothy O'Donnell, Siva Reddy

Language models (LMs) may appear insensitive to word order changes in natural language understanding (NLU) tasks. In this paper, we propose that linguistic redundancy can explain this phenomenon, whereby word order and o…

Natural Language UnderstandingRTESST-2

Constructing Hierarchical Q&A Datasets for Video Story Understanding

2019-04-01 · Yu-Jung Heo, Kyoung-Woon On, SeongHo Choi, Jaeseo Lim 외

Video understanding is emerging as a new paradigm for studying human-like AI. Question-and-Answering (Q&A) is used as a general benchmark to measure the level of intelligence for video understanding. While several previo…

Video Understanding