paper-with-me

홈 › Papers

Non-readily identifiable data collaboration analysis for multiple datasets including personal information

2022-08-31 · Akira Imakura, Tetsuya Sakurai, Yukihiko Okada, Tomoya Fujii, Teppei Sakamoto, Hiroyuki Abe

Multi-source data fusion, in which multiple data sources are jointly analyzed to obtain improved information, has considerable research attention. For the datasets of multiple medical institutions, data confidentiality and cross-institutional communication are critical. In such cases, data collaboration (DC) analysis by sharing dimensionality-reduced intermediate representations without iterative cross-institutional communications may be appropriate. Identifiability of the shared data is essential when analyzing data including personal information. In this study, the identifiability of the DC analysis is investigated. The results reveals that the shared intermediate representations are readily identifiable to the original data for supervised learning. This study then proposes a non-readily identifiable DC analysis only sharing non-readily identifiable data for multiple medical datasets including personal information. The proposed method solves identifiability concerns based on a random sample permutation, the concept of interpretable DC analysis, and usage of functions that cannot be reconstructed. In numerical experiments on medical datasets, the proposed method exhibits a non-readily identifiability while maintaining a high recognition performance of the conventional DC analysis. For a hospital dataset, the proposed method exhibits a nine percentage point improvement regarding the recognition performance over the local analysis that uses only local dataset.

📄 PDF Abstract BibTeX arXiv:2208.14611

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Identifiable Phenotyping using Constrained Non-Negative Matrix Factorization

2016-08-02 · Shalmali Joshi, Suriya Gunasekar, David Sontag, Joydeep Ghosh

This work proposes a new algorithm for automated and simultaneous phenotyping of multiple co-occurring medical conditions, also referred as comorbidities, using clinical notes from the electronic health records (EHRs). A…

Enhancing Code Quality with Generative AI: Boosting Developer Warning Compliance

2025-05-16 · Hansen Chang, Christian DeLozier

Programmers have long ignored warnings, especially those generated by static analysis tools, due to the potential for false-positives. In some cases, warnings may be indicative of larger issues, but programmers may not u…

CTAP: A Web-Based Tool Supporting Automatic Complexity Analysis

2016-12-01 · WS 2016 12 · Xiaobin Chen, Detmar Meurers

Informed by research on readability and language acquisition, computational linguists have developed sophisticated tools for the analysis of linguistic complexity. While some tools are starting to become accessible on th…

Language AcquisitionManagement

Pan-genome Analysis of Angiosperm Plastomes using PGR-TK

2025-04-28 · Manoj P. Samanta

We present a novel approach for taxonomic analysis of chloroplast genomes in angiosperms using the Pan-genome Research Toolkit (PGR-TK). Comparative plots generated by PGR-TK across diverse angiosperm genera reveal a wid…

FirmCORe: A Benchmark for Structured Reasoning about Inter-Firm Collaboration Opportunities

2026-09-15 · Tian Du, Tiantong Wu, Yafei Wang, Mengyu Liu 외 arxiv

Comprehensive structured data on inter-firm relationships is often scarce or inaccessible because many relationships are privately negotiated, selectively disclosed, and fragmented across proprietary databases. This scar…