paper-with-me

홈 › Papers

Exploring Graph Representations of Logical Forms for Language Modeling

2025-05-20 · Michael Sullivan

We make the case for language models over logical forms (LFLMs), arguing that such models are more data-efficient than their textual counterparts. To that end, we introduce the Graph-based Formal-Logical Distributional Semantics (GFoLDS) prototype, a pretrained LM over graph representations of logical forms, as a proof-of-concept of LFLMs. Using GFoLDS, we present strong experimental evidence that LFLMs can leverage the built-in, basic linguistic knowledge inherent in such models to immediately begin learning more complex patterns. On downstream tasks, we show that GFoLDS vastly outperforms textual, transformer LMs pretrained on similar amounts of data, indicating that LFLMs can learn with substantially less data than models over plain text. Furthermore, we show that the performance of this model is likely to scale with additional parameters and pretraining data, suggesting the viability of LFLMs in real-world applications.

📄 PDF Abstract BibTeX arXiv:2505.14523

Code (1)

mjs227/gfolds 공식 구현 pytorch

Tasks

Language ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

Exploring the Semantic Content of Unsupervised Graph Embeddings: An Empirical Study

2018-06-19 · Stephen Bonner, Ibad Kureshi, John Brennan, Georgios Theodoropoulos 외

Graph embeddings have become a key and widely used technique within the field of graph mining, proving to be successful across a broad range of domains including social, citation, transportation and biological. Graph emb…

Graph EmbeddingGraph Mining

LangTopo: Aligning Language Descriptions of Graphs with Tokenized Topological Modeling

2024-06-19 · Zhong Guan, Hongke Zhao, Likang Wu, Ming He 외

Recently, large language models (LLMs) have been widely researched in the field of graph machine learning due to their outstanding abilities in language comprehension and learning. However, the significant gap between na…

Natural Language Understanding

DiaWUG: A Dataset for Diatopic Lexical Semantic Variation in Spanish

2022-06-01 · LREC 2022 6 · Gioia Baldissin, Dominik Schlechtweg, Sabine Schulte im Walde

We provide a novel dataset – DiaWUG – with judgements on diatopic lexical semantic variation for six Spanish variants in Europe and Latin America. In contrast to most previous meaning-based resources and studies on seman…

Exploring How Generative Adversarial Networks Learn Phonological Representations

2023-05-21 · Jingyi Chen, Micha Elsner

This paper explores how Generative Adversarial Networks (GANs) learn representations of phonological phenomena. We analyze how GANs encode contrastive and non-contrastive nasality in French and English vowels by applying…

LMAE4Eth: Generalizable and Robust Ethereum Fraud Detection by Exploring Transaction Semantics and Masked Graph Embedding

2025-09-04 · Yifan Jia, Yanbin Wang, Jianguo Sun, Ye Tian 외 arxiv

Current Ethereum fraud detection methods rely on context-independent, numerical transaction sequences, failing to capture semantic of account transactions. Furthermore, the pervasive homogeneity in Ethereum transaction r…

Self-Supervised LearningContrastive LearningGraph EmbeddingFraud Detection