paper-with-me

홈 › Papers

New Perspective of Interpretability of Deep Neural Networks

2019-09-12 · Masanari Kimura, Masayuki Tanaka

Deep neural networks (DNNs) are known as black-box models. In other words, it is difficult to interpret the internal state of the model. Improving the interpretability of DNNs is one of the hot research topics. However, at present, the definition of interpretability for DNNs is vague, and the question of what is a highly explanatory model is still controversial. To address this issue, we provide the definition of the human predictability of the model, as a part of the interpretability of the DNNs. The human predictability proposed in this paper is defined by easiness to predict the change of the inference when perturbating the model of the DNNs. In addition, we introduce one example of high human-predictable DNNs. We discuss that our definition will help to the research of the interpretability of the DNNs considering various types of applications.

📄 PDF Abstract BibTeX arXiv:1909.07156

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Interpretability 설명 없음

Similar Papers 제목 키워드 기반

On the Interpretability of Part-Prototype Based Classifiers: A Human Centric Analysis

2023-10-10 · Omid Davoodi, Shayan Mohammadizadehsamakosh, Majid Komeili

Part-prototype networks have recently become methods of interest as an interpretable alternative to many of the current black-box image classifiers. However, the interpretability of these methods from the perspective of …

Transformer Key-Value Memories Are Nearly as Interpretable as Sparse Autoencoders

2025-10-25 · Mengyu Ye, Jun Suzuki, Tatsuro Inaba, Tatsuki Kuribayashi arxiv

Recent interpretability work on large language models (LLMs) has been increasingly dominated by a feature-discovery approach with the help of proxy modules. Then, the quality of features learned by, e.g., sparse auto-enc…

Improving Accuracy Without Losing Interpretability: A ML Approach for Time Series Forecasting

2022-12-13 · Yiqi Sun, Zhengxin Shi, Jianshen Zhang, Yongzhi Qi 외

In time series forecasting, decomposition-based algorithms break aggregate data into meaningful components and are therefore appreciated for their particular advantages in interpretability. Recent algorithms often combin…

MarketingTime SeriesTime Series AnalysisTime Series Forecasting

A Dual-Perspective NLG Meta-Evaluation Framework with Automatic Benchmark and Better Interpretability

2025-02-17 · Xinyu Hu, Mingqi Gao, Li Lin, Zhenghan Yu 외

In NLG meta-evaluation, evaluation metrics are typically assessed based on their consistency with humans. However, we identify some limitations in traditional NLG meta-evaluation approaches, such as issues in handling hu…

Towards Ethical Multi-Agent Systems of Large Language Models: A Mechanistic Interpretability Perspective

2025-12-04 · Jae Hee Lee, Anne Lauscher, Stefano V. Albrecht arxiv

Large language models (LLMs) have been widely deployed in various applications, often functioning as autonomous agents that interact with each other in multi-agent systems. While these systems have shown promise in enhan…