Text-Based Joint Prediction of Numeric and Categorical Attributes of Entities in Knowledge Bases
Collaboratively constructed knowledge bases play an important role in information systems, but are essentially always incomplete. Thus, a large number of models has been developed for Knowledge Base Completion, the task of predicting new attributes of entities given partial descriptions of these entities. Virtually all of these models either concentrate on numeric attributes ({\textless}Italy,GDP,2T{\$}{\textgreater}) or they concentrate on categorical attributes ({\textless}Tim Cook,chairman,Apple{\textgreater}). In this paper, we propose a simple feed-forward neural architecture to jointly predict numeric and categorical attributes based on embeddings learned from textual occurrences of the entities in question. Following insights from multi-task learning, our hypothesis is that due to the correlations among attributes of different kinds, joint prediction improves over separate prediction. Our experiments on seven FreeBase domains show that this hypothesis is true of the two attribute types: we find substantial improvements for numeric attributes in the joint model, while performance remains largely unchanged for categorical attributes. Our analysis indicates that this is the case because categorical attributes, many of which describe membership in various classes, provide useful {`}background knowledge{'} for numeric prediction, while this is true to a lesser degree in the inverse direction.
Code (0)
등록된 구현이 없습니다.
Tasks
AttributeKnowledge Base CompletionMulti-Task LearningPredictionSimilar Papers 제목 키워드 기반
Text-Aware Predictive Monitoring of Business Processes
The real-time prediction of business processes using historical event data is an important capability of modern business process monitoring systems. Existing process prediction methods are able to also exploit the data p…
PredictionTabDLM: Free-Form Tabular Data Generation via Joint Numerical-Language Diffusion
Synthetic tabular data generation has attracted growing attention due to its importance for data augmentation, foundation models, and privacy. However, real-world tabular datasets increasingly contain free-form text fiel…
Tabular Data GenerationData AugmentationLogical Reasoning for Task Oriented Dialogue Systems
In recent years, large pretrained models have been used in dialogue systems to improve successful task completion rates. However, lack of reasoning capabilities of dialogue platforms make it difficult to provide relevant…
Logical ReasoningNegationSynthetic Data GenerationTask-Oriented Dialogue SystemsmultivariateGPT: a decoder-only transformer for multivariate categorical and numeric data
Real-world processes often generate data that are a mix of categorical and numeric values that are recorded at irregular and informative intervals. Discrete token-based approaches are limited in numeric representation ca…
DecoderNo imputation without representation
By filling in missing values in datasets, imputation allows these datasets to be used with algorithms that cannot handle missing values by themselves. However, missing values may in principle contribute useful informatio…
AttributeImputationMissing Values