Probing the statistical properties of enriched co-occurrence networks
Recent studies have explored the addition of virtual edges to word co-occurrence networks using word embeddings to enhance graph representations, particularly for short texts. While these enriched networks have demonstrated some success, the impact of incorporating semantic edges into traditional co-occurrence networks remains uncertain. This study investigates two key statistical properties of text-based network models. First, we assess whether network metrics can effectively distinguish between meaningless and meaningful texts. Second, we analyze whether these metrics are more sensitive to syntactic or semantic aspects of the text. Our results show that incorporating virtual edges can have positive and negative effects, depending on the specific network metric. For instance, the informativeness of the average shortest path and closeness centrality improves in short texts, while the clustering coefficient's informativeness decreases as more virtual edges are added. Additionally, we found that including stopwords affects the statistical properties of enriched networks. Our results can serve as a guideline for determining which network metrics are most appropriate for specific applications, depending on the typical text size and the nature of the problem.
Code (0)
등록된 구현이 없습니다.
Tasks
InformativenessWord EmbeddingsSimilar Papers 제목 키워드 기반
Data Enrichment: Multi-task Learning in High Dimension with Theoretical Guarantees
Given samples from a group of related regression tasks, a data-enriched model describes observations by a common and per-group individual parameters. In high-dimensional regime, each parameter has its own structure such …
Multi-Task LearningVocal Bursts Intensity PredictionThe Length and the Width of the Human Brain Circuit Connections are Strongly Correlated
The correlations of several fundamental properties of human brain connections are investigated in a consensus connectome, constructed from 1064 braingraphs, each on 1015 vertices, corresponding to 1015 anatomical brain a…
Function Words as Statistical Cues for Language Learning
What statistical properties might support learning abstract grammatical knowledge from linear input? We address this question by examining the statistical distribution of function words. Function words have been argued t…
Measure-Theoretic Probability of Complex Co-occurrence and E-Integral
Complex high-dimensional co-occurrence data are increasingly popular from a complex system of interacting physical, biological and social processes in discretely indexed modifiable areal units or continuously indexed loc…
Information-theoretic Probing Explains Reliance on Spurious Features
Most current NLP systems are based on a pre-train-then-fine-tune paradigm, in which a large neural network is first trained in a self-supervised way designed to encourage the network to extract broadly-useful linguistic …