paper-with-me

홈 › Papers

We Need Improved Data Curation and Attribution in AI for Scientific Discovery

2025-04-03 · Mara Graziani, Antonio Foncubierta, Dimitrios Christofidellis, Irina Espejo-Morales, Malina Molnar, Marvin Alberts, Matteo Manica, Jannis Born

As the interplay between human-generated and synthetic data evolves, new challenges arise in scientific discovery concerning the integrity of the data and the stability of the models. In this work, we examine the role of synthetic data as opposed to that of real experimental data for scientific research. Our analyses indicate that nearly three-quarters of experimental datasets available on open-access platforms have relatively low adoption rates, opening new opportunities to enhance their discoverability and usability by automated methods. Additionally, we observe an increasing difficulty in distinguishing synthetic from real experimental data. We propose supplementing ongoing efforts in automating synthetic data detection by increasing the focus on watermarking real experimental data, thereby strengthening data traceability and integrity. Our estimates suggest that watermarking even less than half of the real world data generated annually could help sustain model robustness, while promoting a balanced integration of synthetic and human-generated content.

📄 PDF Abstract BibTeX arXiv:2504.02486

Code (0)

등록된 구현이 없습니다.

Tasks

scientific discovery

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

SciRAG: Adaptive, Citation-Aware, and Outline-Guided Retrieval and Synthesis for Scientific Literature

2025-11-18 · Hang Ding, Yilun Zhao, Tiansheng Hu, Manasi Patwardhan 외 arxiv

The accelerating growth of scientific publications has intensified the need for scalable, trustworthy systems to synthesize knowledge across diverse literature. While recent retrieval-augmented generation (RAG) methods h…

Building Agent Harnesses for Scientific Curation from Multimodal Sources

2026-06-19 · Sheng Zhang, Qin Liu, Renqian Luo, Shufang Xie 외 arxiv

Scientific discovery workflows often depend on structured curation from the literature. This is difficult for current agents because the key evidence is scattered across long text, dense tables, and figures, and the fina…

Think Before You Attribute: Improving the Performance of LLMs Attribution Systems

2025-05-19 · João Eduardo Batista, Emil Vatai, Mohamed Wahib

Large Language Models (LLMs) are increasingly applied in various science domains, yet their broader adoption remains constrained by a critical challenge: the lack of trustworthy, verifiable outputs. Current LLMs often ge…

AttributeRAGSentence

Assessment of the Reliablity of a Model's Decision by Generalizing Attribution to the Wavelet Domain

2023-05-24 · Gabriel Kasmi, Laurent Dubus, Yves-Marie Saint Drenan, Philippe Blanc

Neural networks have shown remarkable performance in computer vision, but their deployment in numerous scientific and technical fields is challenging due to their black-box nature. Scientists and practitioners need to ev…

Prospector Heads: Generalized Feature Attribution for Large Models & Data

2024-02-18 · Gautam Machiraju, Alexander Derry, Arjun Desai, Neel Guha 외

Feature attribution, the ability to localize regions of the input data that are relevant for classification, is an important capability for ML models in scientific and biomedical domains. Current methods for feature attr…