paper-with-me

홈 › Papers

Should I Get Involved? On the Privacy Perils of Mining Software Repositories for Research Participants

2022-02-24 · Melina Vidoni, Nicolás E. Díaz Ferreyra

Mining Software Repositories (MSRs) is an evidence-based methodology that cross-links data to uncover actionable information about software systems. Empirical studies in software engineering often leverage MSR techniques as they allow researchers to unveil issues and flaws in software development so as to analyse the different factors contributing to them. Hence, counting on fine-grained information about the repositories and sources being mined (e.g., server names, and contributors' identities) is essential for the reproducibility and transparency of MSR studies. However, this can also introduce threats to participants' privacy as their identities may be linked to flawed/sub-optimal programming practices (e.g., code smells, improper documentation), or vice-versa. Moreover, this can be extensible to close collaborators and community members resulting "guilty by association". This position paper aims to start a discussion about indirect participation in MSRs investigations, the dichotomy of 'privacy vs. utility' regarding sharing non-aggregated data, and its effects on privacy restrictions and ethical considerations for participant involvement.

📄 PDF Abstract BibTeX arXiv:2202.11969

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Opinion Mining and Analysis: A survey

2013-07-12 · Arti Buche, Dr. M. B. Chandak, Akshay Zadgaonkar

The current research is focusing on the area of Opinion Mining also called as sentiment analysis due to sheer volume of opinion rich web resources such as discussion forums, review sites and blogs are available in digita…

Opinion MiningSentiment AnalysisSurvey

Explanation Hacking: The perils of algorithmic recourse

2024-03-22 · Emily Sullivan, Atoosa Kasirzadeh

We argue that the trend toward providing users with feasible and actionable explanations of AI decisions, known as recourse explanations, comes with ethical downsides. Specifically, we argue that recourse explanations fa…

Property-Driven Synthetic Data Engineering for Data-Scarce Software Systems: Reflections from the Breast Cancer Domain

2026-07-07 · Aurora Francesca Zanenga, Andrea Bombarda, Marsha Chechik, Saverio D'Amico 외 arxiv

Modern software systems increasingly depend on data for analysis, prediction, testing, and decision-making. Yet many important domains, including medicine, safety-critical systems, and regulated industries, lack abundant…

Synthetic Data Generation

Computer-Aided Data Mining: Automating a Novel Knowledge Discovery and Data Mining Process Model for Metabolomics

2019-07-09 · Ahmed BaniMustafa, Nigel Hardy

This work presents MeKDDaM-SAGA, computer-aided automation software for implementing a novel knowledge discovery and data mining process model that was designed for performing justifiable, traceable and reproducible meta…

The Perils of Overreaction

2024-05-13 · Konstantin von Beringe, Mark Whitmeyer

In order to study updating rules, we consider the problem of a malevolent principal screening an imperfectly Bayesian agent. We uncover a fundamental dichotomy between underreaction and overreaction to information. If an…