An Information-theoretic On-line Learning Principle for Specialization in Hierarchical Decision-Making Systems
Information-theoretic bounded rationality describes utility-optimizing decision-makers whose limited information-processing capabilities are formalized by information constraints. One of the consequences of bounded rationality is that resource-limited decision-makers can join together to solve decision-making problems that are beyond the capabilities of each individual. Here, we study an information-theoretic principle that drives division of labor and specialization when decision-makers with information constraints are joined together. We devise an on-line learning rule of this principle that learns a partitioning of the problem space such that it can be solved by specialized linear policies. We demonstrate the approach for decision-making problems whose complexity exceeds the capabilities of individual decision-makers, but can be solved by combining the decision-makers optimally. The strength of the model is that it is abstract and principled, yet has direct applications in classification, regression, reinforcement learning and adaptive control.
Code (0)
등록된 구현이 없습니다.
Tasks
Decision MakingReinforcement LearningSimilar Papers 제목 키워드 기반
Systems of bounded rational agents with information-theoretic constraints
Specialization and hierarchical organization are important features of efficient collaboration in economical, artificial, and biological systems. Here, we investigate the hypothesis that both features can be explained by…
Hierarchical Expert Networks for Meta-Learning
The goal of meta-learning is to train a model on a variety of learning tasks, such that it can adapt to new problems within only a few iterations. Here we propose a principled information-theoretic model that optimally p…
image-classificationImage ClassificationMeta-Learningregression+3Specialization in Hierarchical Learning Systems
Joining multiple decision-makers together is a powerful way to obtain more sophisticated decision-making systems, but requires to address the questions of division of labor and specialization. We investigate in how far i…
Decision MakingDensity EstimationMeta-LearningGeometric Metrics for MoE Specialization: From Fisher Information to Early Failure Detection
Expert specialization is fundamental to Mixture-of-Experts (MoE) model success, yet existing metrics (cosine similarity, routing entropy) lack theoretical grounding and yield inconsistent conclusions under reparameteriza…
Hierarchical Mixture-of-Experts with Two-Stage Optimization
Sparse Mixture-of-Experts (MoE) models scale capacity by routing each token to a small subset of experts. However, their routers exhibit a fundamental trade-off: strong load balancing can suppress expert specialization, …