paper-with-me

Papers

Decoder-based Sense Knowledge Distillation

2026-02-25 · Qitong Wang, Mohammed J. Zaki, Georgios Kollias, Vasileios Kalantzis arxiv

Large language models (LLMs) learn contextual embeddings that capture rich semantic information, yet they often overlook structured lexical knowledge such as word senses and relationships. Prior work has shown that incorporating sense dictionaries can improve knowledge distillation for encoder models, but their application to decoder as generative models remains challenging. In this paper, we introduce Decoder-based Sense Knowledge Distillation (DSKD), a framework that integrates lexical resources into the training of decoder-style LLMs without requiring dictionary lookup at inference time. Extensive experiments on diverse benchmarks demonstrate that DSKD significantly enhances knowledge distillation performance for decoders, enabling generative models to inherit structured semantics while maintaining efficient training.

📄 PDF Abstract BibTeX arXiv:2602.22351

Code (0)

등록된 구현이 없습니다.

Tasks

Knowledge Distillation

Similar Papers 제목 키워드 기반

Symbolic Knowledge Distillation: from General Language Models to Commonsense Models

2021-10-14 · NAACL 2022 7 · Peter West, Chandra Bhagavatula, Jack Hessel, Jena D. Hwang 외

The common practice for training commonsense models has gone from-human-to-corpus-to-machine: humans author commonsense knowledge graphs in order to train commonsense models. In this work, we investigate an alternative, …

Knowledge DistillationKnowledge GraphsLanguage ModelingLanguage Modelling+1

Active Object Detection with Knowledge Aggregation and Distillation from Large Models

2024-05-21 · CVPR 2024 1 · Dejie Yang, Yang Liu

Accurately detecting active objects undergoing state changes is essential for comprehending human interactions and facilitating decision-making. The existing methods for active object detection (AOD) primarily rely on vi…

Active Object DetectionDecision MakingDecoderKnowledge Distillation+3

I2D2: Inductive Knowledge Distillation with NeuroLogic and Self-Imitation

2022-12-19 · Chandra Bhagavatula, Jena D. Hwang, Doug Downey, Ronan Le Bras 외

Commonsense capabilities of pre-trained language models dramatically improve with scale, leading many to believe that scale is the only winning recipe. But is it? Here, we investigate an alternative that a priori seems i…

Imitation LearningKnowledge Distillation

Multimodal Commonsense Knowledge Distillation for Visual Question Answering

2024-11-05 · Shuo Yang, Siwen Luo, Soyeon Caren Han

Existing Multimodal Large Language Models (MLLMs) and Visual Language Pretrained Models (VLPMs) have shown remarkable performances in the general Visual Question Answering (VQA). However, these models struggle with VQA q…

Knowledge DistillationQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

WebChild 2.0 : Fine-Grained Commonsense Knowledge Distillation

2017-07-01 · ACL 2017 7 · T, Niket on, Gerard de Melo, Gerhard Weikum
Knowledge DistillationSemantic Parsing