paper-with-me

홈 › Papers

Learning Metadata-Agnostic Representations for Text-to-SQL In-Context Example Selection

2024-10-17 · Chuhong Mai, Ro-ee Tal, Thahir Mohamed

In-context learning (ICL) is a powerful paradigm where large language models (LLMs) benefit from task demonstrations added to the prompt. Yet, selecting optimal demonstrations is not trivial, especially for complex or multi-modal tasks where input and output distributions differ. We hypothesize that forming task-specific representations of the input is key. In this paper, we propose a method to align representations of natural language questions and those of SQL queries in a shared embedding space. Our technique, dubbed MARLO - Metadata-Agnostic Representation Learning for Text-tO-SQL - uses query structure to model querying intent without over-indexing on underlying database metadata (i.e. tables, columns, or domain-specific entities of a database referenced in the question or query). This allows MARLO to select examples that are structurally and semantically relevant for the task rather than examples that are spuriously related to a certain domain or question phrasing. When used to retrieve examples based on question similarity, MARLO shows superior performance compared to generic embedding models (on average +2.9\%pt. in execution accuracy) on the Spider benchmark. It also outperforms the next best method that masks metadata information by +0.8\%pt. in execution accuracy on average, while imposing a significantly lower inference latency.

📄 PDF Abstract BibTeX arXiv:2410.14049

Code (0)

등록된 구현이 없습니다.

Tasks

In-Context LearningQuestion SimilarityRepresentation LearningText to SQLText-To-SQL

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Robust Document Representations using Latent Topics and Metadata

2020-10-23 · Natraj Raman, Armineh Nourbakhsh, Sameena Shah, Manuela Veloso

Task specific fine-tuning of a pre-trained neural language model using a custom softmax output layer is the de facto approach of late when dealing with document classification problems. This technique is not adequate whe…

Document ClassificationLanguage ModelingLanguage Modelling

Enhancing Radiographic Disease Detection with MetaCheX, a Context-Aware Multimodal Model

2025-09-15 · Nathan He, Cody Chen arxiv

Existing deep learning models for chest radiology often neglect patient metadata, limiting diagnostic accuracy and fairness. To bridge this gap, we introduce MetaCheX, a novel multimodal framework that integrates chest X…

Zero-Shot Clinical Acronym Expansion via Latent Meaning Cells

2020-09-29 · Griffin Adams, Mert Ketenci, Shreyas Bhave, Adler Perotte 외

We introduce Latent Meaning Cells, a deep latent variable model which learns contextualized representations of words by combining local lexical context and metadata. Metadata can refer to granular context, such as sectio…

Representation Learning

Contextualizing ASR Lattice Rescoring with Hybrid Pointer Network Language Model

2020-05-15 · Da-Rong Liu, Chunxi Liu, Frank Zhang, Gabriel Synnaeve 외

Videos uploaded on social media are often accompanied with textual descriptions. In building automatic speech recognition (ASR) systems for videos, we can exploit the contextual information provided by such video metadat…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+2

Multi-Task Reinforcement Learning with Context-based Representations

2021-02-11 · Shagun Sodhani, Amy Zhang, Joelle Pineau

The benefit of multi-task learning over single-task learning relies on the ability to use relations across tasks to improve performance on any single task. While sharing representations is an important mechanism to share…

Multi-Task Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)+1