paper-with-me

홈 › Papers

Modularizing while Training: A New Paradigm for Modularizing DNN Models

2023-06-15 · Binhang Qi, Hailong Sun, Hongyu Zhang, Ruobing Zhao, Xiang Gao

Deep neural network (DNN) models have become increasingly crucial components in intelligent software systems. However, training a DNN model is typically expensive in terms of both time and money. To address this issue, researchers have recently focused on reusing existing DNN models - borrowing the idea of code reuse in software engineering. However, reusing an entire model could cause extra overhead or inherits the weakness from the undesired functionalities. Hence, existing work proposes to decompose an already trained model into modules, i.e., modularizing-after-training, and enable module reuse. Since trained models are not built for modularization, modularizing-after-training incurs huge overhead and model accuracy loss. In this paper, we propose a novel approach that incorporates modularization into the model training process, i.e., modularizing-while-training (MwT). We train a model to be structurally modular through two loss functions that optimize intra-module cohesion and inter-module coupling. We have implemented the proposed approach for modularizing Convolutional Neural Network (CNN) models in this work. The evaluation results on representative models demonstrate that MwT outperforms the state-of-the-art approach. Specifically, the accuracy loss caused by MwT is only 1.13 percentage points, which is 1.76 percentage points less than that of the baseline. The kernel retention rate of the modules generated by MwT is only 14.58%, with a reduction of 74.31% over the state-of-the-art approach. Furthermore, the total time cost required for training and modularizing is only 108 minutes, half of the baseline.

📄 PDF Abstract BibTeX arXiv:2306.09376

Code (1)

qibinhang/mwt 공식 구현 pytorch

Similar Papers 제목 키워드 기반

NeMo: A Neuron-Level Modularizing-While-Training Approach for Decomposing DNN Models

2025-08-15 · Xiaohan Bi, Binhang Qi, Hailong Sun, Xiang Gao 외 arxiv

With the growing incorporation of deep neural network (DNN) models into modern software systems, the prohibitive construction costs have become a significant challenge. Model reuse has been widely applied to reduce train…

Contrastive Learning

Modularizing Educational LLM-Agency for Fostering Responsible Learning Assistance

2026-05-28 · Julius Gabelmann, Felix Jahn, Kevin Baum, Sophie van Rossum 외 arxiv

The widespread adoption of AI chatbots in education will drastically change learning, making responsible deployment a critical concern. While large language models (LLMs) might have access to sources discussing insights …

Exploring Domain Robust Lightweight Reward Models based on Router Mechanism

2024-07-24 · Hyuk Namgoong, Jeesu Jung, SangKeun Jung, YoonHyung Roh

Recent advancements in large language models have heavily relied on the large reward model from reinforcement learning from human feedback for fine-tuning. However, the use of a single reward model across various domains…

Language ModelingLanguage ModellingMixture-of-ExpertsSmall Language Model

MUSE: Modularizing Unsupervised Sense Embeddings

2017-04-15 · EMNLP 2017 9 · Guang-He Lee, Yun-Nung Chen

This paper proposes to address the word sense ambiguity issue in an unsupervised manner, where word sense representations are learned along a word sense selection mechanism given contexts. Prior work focused on designing…

Reinforcement LearningReinforcement Learning (RL)Representation Learning

RobotFleet: An Open-Source Framework for Centralized Multi-Robot Task Planning

2025-10-12 · Rohan Gupta, Trevor Asbery, Zain Merchant, Abrar Anwar 외 arxiv

Coordinating heterogeneous robot fleets to achieve multiple goals is challenging in multi-robot systems. We introduce an open-source and extensible framework for centralized multi-robot task planning and scheduling that …

Robot Task Planning