MEME: Generating RNN Model Explanations via Model Extraction
Recurrent Neural Networks (RNNs) have achieved remarkable performance on a range of tasks. A key step to further empowering RNN-based approaches is improving their explainability and interpretability. In this work we present MEME: a model extraction approach capable of approximating RNNs with interpretable models represented by human-understandable concepts and their interactions. We demonstrate how MEME can be applied to two multivariate, continuous data case studies: Room Occupation Prediction, and In-Hospital Mortality Prediction. Using these case-studies, we show how our extracted models can be used to interpret RNNs both locally and globally, by approximating RNN decision-making via interpretable concept interactions.
Code (1)
Tasks
Decision MakingmodelModel extractionMortality PredictionOccupation predictionPredictionSimilar Papers 제목 키워드 기반
MEME: Generating RNN Model Explanations via Model Extraction
Recurrent Neural Networks (RNNs) have achieved remarkable performance on a range of tasks. A key step to further empowering RNN-based approaches is improving their explainability and interpretability. In this work we pre…
Decision MakingmodelModel extractionMortality Prediction+2What do you MEME? Generating Explanations for Visual Semantic Role Labelling in Memes
Memes are powerful means for effective communication on social media. Their effortless amalgamation of viral visuals and compelling messages can have far-reaching implications with proper marketing. Previous research on …
MarketingMulti-Task LearningSemantic Role LabelingText GenerationMeme-ingful Analysis: Enhanced Understanding of Cyberbullying in Memes Through Multimodal Explanations
Internet memes have gained significant influence in communicating political, psychological, and sociocultural ideas. While memes are often humorous, there has been a rise in the use of memes for trolling and cyberbullyin…
Adapting Reinforcement Learning with Chain-of-Thought Supervision for Explainable Detection of Hateful and Propagandistic Memes
Hateful and propagandistic memes exploit the interplay between images and text to convey harmful intent that neither modality reveals alone. Although thinking-based multimodal large language models (MLLMs) have advanced …
Reinforcement LearningDecoding the Underlying Meaning of Multimodal Hateful Memes
Recent studies have proposed models that yielded promising performance for the hateful meme classification task. Nevertheless, these proposed models do not generate interpretable explanations that uncover the underlying …
BenchmarkingHateful Meme ClassificationMeme Classification