Paying down metadata debt: learning the representation of concepts using topic models
We introduce a data management problem called metadata debt, to identify the mapping between data concepts and their logical representations. We describe how this mapping can be learned using semisupervised topic models based on low-rank matrix factorizations that account for missing and noisy labels, coupled with sparsity penalties to improve localization and interpretability. We introduce a gauge transformation approach that allows us to construct explicit associations between topics and concept labels, and thus assign meaning to topics. We also show how to use this topic model for semisupervised learning tasks like extrapolating from known labels, evaluating possible errors in existing labels, and predicting missing features. We show results from this topic model in predicting subject tags on over 25,000 datasets from Kaggle.com, demonstrating the ability to learn semantically meaningful features.
Code (0)
등록된 구현이 없습니다.
Tasks
ManagementTopic ModelsSimilar Papers 제목 키워드 기반
Capital Structure in U.S., a Quantile Regression Approach with Macroeconomic Impacts
The major perspective of this paper is to provide more evidence into the empirical determinants of capital structure adjustment in different macroeconomics states by focusing and discussing the relative importance of fir…
quantile regressionregressionOn The Choice Between A Sale-Leaseback And Debt
This article introduces decision models for commercial real estate leasing. The concepts and models developed in the article can also be applied to equipment leasing and other types of leasing.
Loss Rate Forecasting Framework Based on Macroeconomic Changes: Application to US Credit Card Industry
A major part of the balance sheets of the largest US banks consists of credit card portfolios. Hence, managing the charge-off rates is a vital task for the profitability of the credit card industry. Different macroeconom…
Brain-DiT: A Universal Multi-state fMRI Foundation Model with Metadata-Conditioned Pretraining
Current fMRI foundation models primarily rely on a limited range of brain states and mismatched pretraining tasks, restricting their ability to learn generalized representations across diverse brain states. We present Br…
Metadata-Aligned 3D MRI Representations for Contrast Understanding and Quality Control
Magnetic Resonance Imaging suffers from substantial data heterogeneity and the absence of standardized contrast labels across scanners, protocols, and institutions, which severely limits large-scale automated analysis. A…