paper-with-me

홈 › Papers

Struc-EMB: The Potential of Structure-Aware Encoding in Language Embeddings

2025-10-09 · Shikun Liu, Haoyu Wang, Mufei Li, Pan Li arxiv

Text embeddings from Large Language Models (LLMs) have become foundational for numerous applications. However, these models typically operate on raw text, overlooking the rich structural information, such as hyperlinks or citations, that provides crucial context in many real-world datasets. This paper introduces and systematically evaluates a new paradigm for generating structure-aware text embeddings by integrating these structural relations directly into the LLM's internal encoding process, rather than relying on traditional post-hoc aggregation. We investigate two primary in-process methods: sequential concatenation and parallel caching. Through extensive zero-shot experiments across retrieval, clustering, classification, and recommendation tasks, we demonstrate that our structure-aware approaches consistently outperform both text-only and post-hoc baselines. Our analysis reveals critical trade-offs: sequential concatenation excels with noisy, moderate-length contexts, while parallel caching scales more effectively to long, high-signal contexts but is more susceptible to distractors. To address the challenge of noisy structural data, we also introduce and validate two effective techniques: Context Distillation and Semantic Balancing. This work provides the first comprehensive analysis of in-process structure-aware encoding, offering a blueprint for building more powerful and contextually aware embedding models.

📄 PDF Abstract BibTeX arXiv:2510.08774

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Subgraph-Aware Training of Language Models for Knowledge Graph Completion Using Structure-Aware Contrastive Learning

2024-07-17 · Youmin Ko, Hyemin Yang, Taeuk Kim, Hyunjoon Kim

Fine-tuning pre-trained language models (PLMs) has recently shown a potential to improve knowledge graph completion (KGC). However, most PLM-based methods focus solely on encoding textual information, neglecting the long…

Contrastive LearningInductive BiasKnowledge Graph CompletionKnowledge Graphs

SADGA: Structure-Aware Dual Graph Aggregation Network for Text-to-SQL

2021-11-01 · NeurIPS 2021 12 · Ruichu Cai, Jinjie Yuan, Boyan Xu, Zhifeng Hao

The Text-to-SQL task, aiming to translate the natural language of the questions into SQL queries, has drawn much attention recently. One of the most challenging problems of Text-to-SQL is how to generalize the trained mo…

Semantic ParsingText to SQLText-To-SQL

CATE: Computation-aware Neural Architecture Encoding with Transformers

2021-02-14 · Shen Yan, Kaiqiang Song, Fei Liu, Mi Zhang

Recent works (White et al., 2020a; Yan et al., 2020) demonstrate the importance of architecture encodings in Neural Architecture Search (NAS). These encodings encode either structure or computation information of the neu…

AutoMLNeural Architecture SearchRepresentation LearningUnsupervised Pre-training

Structural Adapters in Pretrained Language Models for AMR-to-text Generation

2021-03-16 · EMNLP 2021 11 · Leonardo F. R. Ribeiro, Yue Zhang, Iryna Gurevych

Pretrained language models (PLM) have recently advanced graph-to-text generation, where the input graph is linearized into a sequence and fed into the PLM to obtain its representation. However, efficiently encoding the g…

AMR-to-Text GenerationData-to-Text GenerationText Generation

XF2T: Cross-lingual Fact-to-Text Generation for Low-Resource Languages

2022-09-22 · Shivprasad Sagare, Tushar Abhishek, Bhavyajeet Singh, Anubhav Sharma 외

Multiple business scenarios require an automated generation of descriptive human-readable text from structured input data. Hence, fact-to-text generation systems have been developed for various downstream tasks like gene…

Data-to-Text GenerationDescriptiveQuestion AnsweringText Generation