paper-with-me

Papers

Language Model Representations for Efficient Few-Shot Tabular Classification

2026-01-21 · Inwon Kang, Parikshit Ram, Yi Zhou, Horst Samulowitz, Oshani Seneviratne arxiv

The Web is a rich source of structured data in the form of tables, from product catalogs and knowledge bases to scientific datasets. However, the heterogeneity of the structure and semantics of these tables makes it challenging to build a unified method that can effectively leverage the information they contain. Meanwhile, Large language models (LLMs) are becoming an increasingly integral component of web infrastructure for tasks like semantic search. This raises a crucial question: can we leverage these already-deployed LLMs to classify structured data in web-native tables (e.g., product catalogs, knowledge base exports, scientific data portals), avoiding the need for specialized models or extensive retraining? This work investigates a lightweight paradigm, $\textbf{Ta}$ble $\textbf{R}$epresentation with $\textbf{L}$anguage Model~($\textbf{TaRL}$), for few-shot tabular classification that directly utilizes semantic embeddings of individual table rows. We first show that naive application of these embeddings underperforms compared to specialized tabular models. We then demonstrate that their potentials can be unlocked with two key techniques: removing the common component from all embeddings and calibrating the softmax temperature. We show that a simple meta-learner, trained on handcrafted features, can learn to predict an appropriate temperature. This approach achieves performance comparable to state-of-the-art models in low-data regimes ($k \leq 32$) of semantically-rich tables. Our findings demonstrate the viability of reusing existing LLM infrastructure for efficient semantics-driven pathway to reuse existing LLM infrastructure for Web table understanding.

📄 PDF Abstract BibTeX arXiv:2602.15844

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

TabLLM: Few-shot Classification of Tabular Data with Large Language Models

2022-10-19 · Stefan Hegselmann, Alejandro Buendia, Hunter Lang, Monica Agrawal 외

We study the application of large language models to zero-shot and few-shot classification of tabular data. We prompt the large language model with a serialization of the tabular data to a natural-language string, togeth…

ClassificationDeep LearningLanguage ModelingLanguage Modelling+4

Multilingual Cognitive Impairment Detection in the Era of Foundation Models

2026-04-08 · Damar Hoogland, Boshko Koloski, Jaya Caporusso, Tine Kolenik 외 arxiv

We evaluate cognitive impairment (CI) classification from transcripts of speech in English, Slovene, and Korean. We compare zero-shot large language models (LLMs) used as direct classifiers under three input settings -- …

LLMTabBench: Evaluating LLMs on Binary Tabular Classification From Zero to Few Shots

2026-05-23 · Daria Grushina, Kseniia Kuvshinova, Alina Kostromina, Aziz Temirkhanov 외 arxiv

Supervised classification on tabular data remains a central machine learning task, but its dependence on large labeled datasets limits its applicability in data-scarce settings. Few-shot methods such as TabPFN achieve st…

Revisiting CLIP: Efficient Alignment of 3D MRI and Tabular Data using Domain-Specific Foundation Models

2025-01-23 · Jakob Krogh Petersen, Valdemar Licht, Mads Nielsen, Asbjørn Munk

Multi-modal models require aligned, shared embedding spaces. However, common CLIP-based approaches need large amounts of samples and do not natively support 3D or tabular data, both of which are crucial in the medical do…

Image RetrievalRetrievalzero-shot-classificationZero-shot Image Retrieval+1

Large Scale Transfer Learning for Tabular Data via Language Modeling

2024-06-17 · Josh Gardner, Juan C. Perdomo, Ludwig Schmidt

Tabular data -- structured, heterogeneous, spreadsheet-style data with rows and columns -- is widely used in practice across many domains. However, while recent foundation models have reduced the need for developing task…

Language ModelingLanguage ModellingLarge Language ModelPrediction+1