Papers Faithfulness Critic
“Faithfulness Critic” 태그가 달린 논문 1편 · 필터 해제
Are self-explanations from Large Language Models faithful?
2024-01-15
· Andreas Madsen, Sarath Chandar, Siva Reddy
Instruction-tuned Large Language Models (LLMs) excel at many tasks and will even explain their reasoning, so-called self-explanations. However, convincing and wrong self-explanations can lead to unsupported confidence in…
counterfactualFaithfulness CriticNatural Language InferenceQuestion Answering+2
1–1 / 1