paper-with-me

홈 › Papers

Comparing Natural and Synthetic Structured Data: A Study of the Passive Verb Alternation in French and Italian

2026-03-26 · Giuseppe Samo, Paola Merlo arxiv

This study compares the impact of natural and synthetic data on training and evaluating large language models (LLMs), using the case of passive verb alternation in French and Italian. We use Blackbird Language Matrices (BLMs), structured datasets designed to probe linguistic knowledge of underlying patterns across sentence sets. We compare structured templates instantiated with natural sentences extracted from Universal Dependencies to structured templates of synthetic sentences. Experiments show that while models achieve ceiling performance when trained and tested on synthetic datasets, they do not reliably generalize to natural sentences. In contrast, models trained on natural data exhibit robust performance across both natural and synthetic test suites, demonstrating their superior ability to capture abstract linguistic patterns. These results corroborate the value of natural data and of structured set ups in linguistic evaluation for probing LLMs' syntactic and semantic knowledge.

📄 PDF Abstract BibTeX arXiv:2603.25227

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Struct-Bench: A Benchmark for Differentially Private Structured Text Generation

2025-09-12 · Shuaiqi Wang, Vikas Raunak, Arturs Backurs, Victor Reis 외 arxiv

Differentially private (DP) synthetic data generation is a promising technique for utilizing private datasets that otherwise cannot be exposed for model training or other analytics. While much research literature has foc…

Synthetic Data GenerationSynthetic Data EvaluationText Generation

STRUCTURED ALIGNMENT NETWORKS

2018-01-01 · ICLR 2018 1 · Yang Liu, Matt Gardner

Many tasks in natural language processing involve comparing two sentences to compute some notion of relevance, entailment, or similarity. Typically this comparison is done either at the word level or at the sentence lev…

Natural Language InferenceSentence

Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls

2025-10-02 · Feiyang Kang, Newsha Ardalani, Michael Kuchnik, Youssef Emad 외 arxiv

Training data plays a crucial role in Large Language Models (LLM) scaling, yet high quality data is of limited supply. Synthetic data techniques offer a potential path toward sidestepping these limitations. We conduct a …

Enhancing Structured-Data Retrieval with GraphRAG: Soccer Data Case Study

2024-09-26 · Zahra Sepasdar, Sushant Gautam, Cise Midoglu, Michael A. Riegler 외

Extracting meaningful insights from large and complex datasets poses significant challenges, particularly in ensuring the accuracy and relevance of retrieved information. Traditional data retrieval methods such as sequen…

Information RetrievalKnowledge GraphsLanguage ModelingLanguage Modelling+3

A Short Note on Event-Study Synthetic Difference-in-Differences Estimators

2024-07-05 · Diego Ciccia

I propose an event study extension of Synthetic Difference-in-Differences (SDID) estimators. I show that, in simple and staggered adoption designs, estimators from Arkhangelsky et al. (2021) can be disaggregated into dyn…