paper-with-me

Papers

Does Feasibility Matter? Understanding the Impact of Feasibility on Synthetic Training Data

2025-05-15 · YiWen Liu, Jessica Bader, Jae Myung Kim

With the development of photorealistic diffusion models, models trained in part or fully on synthetic data achieve progressively better results. However, diffusion models still routinely generate images that would not exist in reality, such as a dog floating above the ground or with unrealistic texture artifacts. We define the concept of feasibility as whether attributes in a synthetic image could realistically exist in the real-world domain; synthetic images containing attributes that violate this criterion are considered infeasible. Intuitively, infeasible images are typically considered out-of-distribution; thus, training on such images is expected to hinder a model's ability to generalize to real-world data, and they should therefore be excluded from the training set whenever possible. However, does feasibility really matter? In this paper, we investigate whether enforcing feasibility is necessary when generating synthetic training data for CLIP-based classifiers, focusing on three target attributes: background, color, and texture. We introduce VariReal, a pipeline that minimally edits a given source image to include feasible or infeasible attributes given by the textual prompt generated by a large language model. Our experiments show that feasibility minimally affects LoRA-fine-tuned CLIP performance, with mostly less than 0.3% difference in top-1 accuracy across three fine-grained datasets. Also, the attribute matters on whether the feasible/infeasible images adversarially influence the classification performance. Finally, mixing feasible and infeasible images in training datasets does not significantly impact performance compared to using purely feasible or infeasible datasets.

📄 PDF Abstract BibTeX arXiv:2505.10551

Code (1)

yiveen/syntheticdatafeasibility 공식 구현 pytorch

Tasks

AttributeLarge Language Model

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
SET Dynamic Sparse Training method where weight mask is updated randomly periodically
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Matter-of-Fact: A Benchmark for Verifying the Feasibility of Literature-Supported Claims in Materials Science

2025-06-04 · Peter Jansen, Samiah Hassan, Ruoyao Wang

Contemporary approaches to assisted scientific discovery use language models to automatically generate large numbers of potential hypothesis to test, while also automatically generating code-based experiments to test tho…

ArticlesCode GenerationRetrieval-augmented Generationscientific discovery

Machine learning 2.0 : Engineering Data Driven AI Products

2018-07-01 · James Max Kanter, Benjamin Schreck, Kalyan Veeramachaneni

ML 2.0: In this paper, we propose a paradigm shift from the current practice of creating machine learning models - which requires months-long discovery, exploration and "feasibility report" generation, followed by re-eng…

BIG-bench Machine Learning

SFBench: The SciFy Scientific Feasibility Benchmark

2026-06-28 · Cash Costello, James Mayfield, Elsbeth Turcan, Christine Piatko 외 arxiv

We present SFBench, a benchmark dataset for evaluating systems that assess the feasibility of scientific claims. SFBench includes 197 claims in materials science, each annotated with a ground-truth feasibility score on a…

Feasibility and stability in large Lotka Volterra systems with interaction structure

2022-11-23 · Xiaoyuan Liu, George W. A. Constable, Jonathan W. Pitchford

Complex system stability can be studied via linear stability analysis using Random Matrix Theory (RMT) or via feasibility (requiring positive equilibrium abundances). Both approaches highlight the importance of interacti…

Computational Feasibility of Clustering under Clusterability Assumptions

2015-01-02 · Shai Ben-David

It is well known that most of the common clustering objectives are NP-hard to optimize. In practice, however, clustering is being routinely carried out. One approach for providing theoretical understanding of this seemin…

Clustering