paper-with-me

홈 › Papers

Peer-Ranked Precision: Creating a Foundational Dataset for Fine-Tuning Vision Models from DataSeeds' Annotated Imagery

2025-06-06 · Sajjad Abdoli, Freeman Lewin, Gediminas Vasiliauskas, Fabian Schonholz

The development of modern Artificial Intelligence (AI) models, particularly diffusion-based models employed in computer vision and image generation tasks, is undergoing a paradigmatic shift in development methodologies. Traditionally dominated by a "Model Centric" approach, in which performance gains were primarily pursued through increasingly complex model architectures and hyperparameter optimization, the field is now recognizing a more nuanced "Data-Centric" approach. This emergent framework foregrounds the quality, structure, and relevance of training data as the principal driver of model performance. To operationalize this paradigm shift, we introduce the DataSeeds.AI sample dataset (the "DSD"), initially comprised of approximately 10,610 high-quality human peer-ranked photography images accompanied by extensive multi-tier annotations. The DSD is a foundational computer vision dataset designed to usher in a new standard for commercial image datasets. Representing a small fraction of DataSeed.AI's 100 million-plus image catalog, the DSD provides a scalable foundation necessary for robust commercial and multimodal AI development. Through this in-depth exploratory analysis, we document the quantitative improvements generated by the DSD on specific models against known benchmarks and make the code and the trained models used in our evaluation publicly available.

📄 PDF Abstract BibTeX arXiv:2506.05673

Code (1)

DataSeeds-ai/DSD-finetune-blip-llava pytorch

Tasks

Hyperparameter OptimizationImage Generation

Similar Papers 제목 키워드 기반

Language Fairness in Multilingual Information Retrieval

2024-05-02 · Eugene Yang, Thomas Jänich, James Mayfield, Dawn Lawrie

Multilingual information retrieval (MLIR) considers the problem of ranking documents in several languages for a query expressed in a language that may differ from any of those languages. Recent work has observed that app…

FairnessInformation RetrievalRetrieval

How to Find Fantastic AI Papers: Self-Rankings as a Powerful Predictor of Scientific Impact Beyond Peer Review

2025-10-02 · Buxin Su, Natalie Collina, Garrett Wen, Didong Li 외 arxiv

Peer review in academic research aims not only to ensure factual correctness but also to identify work of high scientific potential that can shape future research directions. This task is especially critical in fast-movi…

Peer-Voted LLM-Agent Stress Tests Find Feed-Induced Lexical Convergence but No Reliable Matched-Exposure Advantage for Distributed Sources

2026-08-20 · Rana Muhammad Usman, Dominic Williamson hf

Population-level behavior in large-language-model (LLM) agents cannot be characterized by single-agent benchmarks. We introduce PV-SST, a peer-voted social-platform testbed, and report a separately frozen, preregistered …

Identity Theft in AI Conference Peer Review

2025-08-06 · Nihar B. Shah, Melisa Bok, Xukun Liu, Andrew McCallum arxiv

We discuss newly uncovered cases of identity theft in the scientific peer-review process within artificial intelligence (AI) research, with broader implications for other academic procedures. We detail how dishonest rese…

What Can We Do to Improve Peer Review in NLP?

2020-10-08 · Findings of the Association for Computational Linguistics 2020 · Anna Rogers, Isabelle Augenstein

Peer review is our best tool for judging the quality of conference submissions, but it is becoming increasingly spurious. We argue that a part of the problem is that the reviewers and area chairs face a poorly defined ta…