paper-with-me

Papers

The Reverse Turing Test for Evaluating Interpretability Methods on Unknown Tasks

2020-10-15 · NeurIPS Workshop HAMLETS 2020 12 · Anonymous

The Turing Test evaluates a computer program’s ability to mimic human behaviour. The Reverse Turing Test, reversely, evaluates a human’s ability to mimic machine behaviour in a forward prediction task. We propose to use the Reverse Turing Test to evaluate the quality of interpretability methods. The Reverse Turing Test improves on previous experimental protocols for human evaluation of interpretability methods by a) including a training phase, and b) masking the task, which, combined, enables us to evaluate models independently of their quality, in a way that is unbiased by the participants' previous exposure to the task. We present a human evaluation of LIME across five NLP tasks in a Latin Square design and analyze the effect of masking the task in forward prediction experiments. Additionally, we demonstrate a fundamental limitation of LIME and show how this limitation is detrimental for human forward prediction in some NLP tasks.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Prediction

Similar Papers 제목 키워드 기반

RETAIN: An Interpretable Predictive Model for Healthcare using Reverse Time Attention Mechanism

2016-08-19 · NeurIPS 2016 12 · Edward Choi, Mohammad Taha Bahadori, Joshua A. Kulas, Andy Schuetz 외

Accuracy and interpretability are two dominant features of successful predictive models. Typically, a choice must be made in favor of complex black box models such as recurrent neural networks (RNN) for accuracy versus l…

Disease Trajectory Forecasting

User Friendly Line CAPTCHAs

2014-02-04 · A. K. B. Karunathilake, B. M. D. Balasuriya, R. G. Ragel

CAPTCHAs or reverse Turing tests are real-time assessments used by programs (or computers) to tell humans and machines apart. This is achieved by assigning and assessing hard AI problems that could only be solved easily …

A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models

2024-07-02 · Daking Rai, Yilun Zhou, Shi Feng, Abulhair Saparov 외

Mechanistic interpretability (MI) is an emerging sub-field of interpretability that seeks to understand a neural network model by reverse-engineering its internal computations. Recently, MI has garnered significant atten…

Navigate

Distilled Reverse Attention Network for Open-world Compositional Zero-Shot Learning

2023-03-01 · ICCV 2023 1 · Yun Li, Zhe Liu, Saurav Jha, Sally Cripps 외

Open-World Compositional Zero-Shot Learning (OW-CZSL) aims to recognize new compositions of seen attributes and objects. In OW-CZSL, methods built on the conventional closed-world setting degrade severely due to the unco…

Compositional Zero-Shot LearningKnowledge DistillationZero-Shot Learning

Deceiving computers in Reverse Turing Test through Deep Learning

2020-06-01 · Jimut Bahan Pal

It is increasingly becoming difficult for human beings to work on their day to day life without going through the process of reverse Turing test, where the Computers tests the users to be humans or not. Almost every webs…

CAPTCHA DetectionDeep Learning