paper-with-me

홈 › Papers

Towards Trustworthy Dataset Distillation

2023-07-18 · Shijie Ma, Fei Zhu, Zhen Cheng, Xu-Yao Zhang

Efficiency and trustworthiness are two eternal pursuits when applying deep learning in real-world applications. With regard to efficiency, dataset distillation (DD) endeavors to reduce training costs by distilling the large dataset into a tiny synthetic dataset. However, existing methods merely concentrate on in-distribution (InD) classification in a closed-world setting, disregarding out-of-distribution (OOD) samples. On the other hand, OOD detection aims to enhance models' trustworthiness, which is always inefficiently achieved in full-data settings. For the first time, we simultaneously consider both issues and propose a novel paradigm called Trustworthy Dataset Distillation (TrustDD). By distilling both InD samples and outliers, the condensed datasets are capable of training models competent in both InD classification and OOD detection. To alleviate the requirement of real outlier data, we further propose to corrupt InD samples to generate pseudo-outliers, namely Pseudo-Outlier Exposure (POE). Comprehensive experiments on various settings demonstrate the effectiveness of TrustDD, and POE surpasses the state-of-the-art method Outlier Exposure (OE). Compared with the preceding DD, TrustDD is more trustworthy and applicable to open-world scenarios. Our code is available at https://github.com/mashijie1028/TrustDD

📄 PDF Abstract BibTeX arXiv:2307.09165

Code (2)

mashijie1028/trustdd 공식 구현 pytorch
Guang000/Awesome-Dataset-Distillation

Tasks

Dataset Distillation

Similar Papers 제목 키워드 기반

Task-Driven Causal Feature Distillation: Towards Trustworthy Risk Prediction

2023-12-20 · Zhixuan Chu, Mengxuan Hu, Qing Cui, Longfei Li 외

Since artificial intelligence has seen tremendous recent successes in many areas, it has sparked great interest in its potential for trustworthy and interpretable risk prediction. However, most models lack causal reasoni…

Prediction

FIRST: Teach A Reliable Large Language Model Through Efficient Trustworthy Distillation

2024-08-22 · Kashun Shum, Minrui Xu, Jianshu Zhang, Zixin Chen 외

Large language models (LLMs) have become increasingly prevalent in our daily lives, leading to an expectation for LLMs to be trustworthy -- - both accurate and well-calibrated (the prediction confidence should align with…

Language ModelingLanguage ModellingLarge Language Model

TT-XAI: Trustworthy Clinical Text Explanations via Keyword Distillation and LLM Reasoning

2025-07-30 · Kristian Miok, Blaz Škrlj, Daniela Zaharie, Marko Robnik Šikonja arxiv

Clinical language models often struggle to provide trustworthy predictions and explanations when applied to lengthy, unstructured electronic health records (EHRs). This work introduces TT-XAI, a lightweight and effective…

Trust-Aware Diversion for Data-Effective Distillation

2025-02-07 · Zhuojie Wu, Yanbin Liu, Xin Shen, Xiaofeng Cao 외

Dataset distillation compresses a large dataset into a small synthetic subset that retains essential information. Existing methods assume that all samples are perfectly labeled, limiting their real-world applications whe…

Dataset DistillationModel Optimization

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence

2026-05-07 · Xinquan Chen, Zhenyun Yin, Shan He, Bin Huang 외 arxiv

As large models evolve from conversational assistants into autonomous agents, challenges increasingly arise from long-horizon decision making, tool use, and real environment interaction. Existing agenticinfrastructure re…

Reinforcement LearningDecision Making