paper-with-me

홈 › Papers

Privacy-Preserving Clustering of Unstructured Big Data for Cloud-Based Enterprise Search Solutions

2020-05-22 · SM Zobaed, Mohsen Amini Salehi

Cloud-based enterprise search services (e.g., Amazon Kendra) are enchanting to big data owners by providing them with convenient search solutions over their enterprise big datasets. However, individuals and businesses that deal with confidential big data (eg, credential documents) are reluctant to fully embrace such services, due to valid concerns about data privacy. Solutions based on client-side encryption have been explored to mitigate privacy concerns. Nonetheless, such solutions hinder data processing, specifically clustering, which is pivotal in dealing with different forms of big data. For instance, clustering is critical to limit the search space and perform real-time search operations on big datasets. To overcome the hindrance in clustering encrypted big data, we propose privacy-preserving clustering schemes for three forms of unstructured encrypted big datasets, namely static, semi-dynamic, and dynamic datasets. To preserve data privacy, the proposed clustering schemes function based on statistical characteristics of the data and determine (A) the suitable number of clusters and (B) appropriate content for each cluster. Experimental results obtained from evaluating the clustering schemes on three different datasets demonstrate between 30% to 60% improvement on the clusters' coherency compared to other clustering schemes for encrypted data. Employing the clustering schemes in a privacy-preserving enterprise search system decreases its search time by up to 78%, while increases the search accuracy by up to 35%.

📄 PDF Abstract BibTeX arXiv:2005.11317

Code (0)

등록된 구현이 없습니다.

Tasks

ClusteringPrivacy Preservingvalid

Similar Papers 제목 키워드 기반

Privacy-Preserving Clustering: A New ApproachBased on Invariant Order Encryption

2020-12-20 · Mihail-Iulian Pleșa, Cezar Pleșca

Cloud computing is increasingly used. One main use of cloud computing is the running of a machine learning algorithm. Due to the large amount of data required for these algorithms, they can …

Cloud ComputingClusteringPrivacy Preserving

Towards Federated Clustering: A Client-wise Private Graph Aggregation Framework

2025-11-14 · Guanxiong He, Jie Wang, Liaoyuan Tang, Zheng Wang 외 arxiv

Federated clustering addresses the critical challenge of extracting patterns from decentralized, unlabeled data. However, it is hampered by the flaw that current approaches are forced to accept a compromise between perfo…

Graph Clustering

HyFedRAG: A Federated Retrieval-Augmented Generation Framework for Heterogeneous and Privacy-Sensitive Data

2025-09-08 · Cheng Qian, Hainan Zhang, Yongxin Tong, Hong-Wei Zheng 외 arxiv

Centralized RAG pipelines struggle with heterogeneous and privacy-sensitive data, especially in distributed healthcare settings where patient data spans SQL, knowledge graphs, and clinical notes. Clinicians face difficul…

Knowledge Graphs

A Privacy Preserving Data Publishing Middleware for Unstructured, Textual Social Media Data

2020-05-01 · LREC 2020 5 · Prasadi Abeywardana, Uthayasanker Thayasivam

Privacy is going to be an integral part of data science and analytics in the coming years. The next hype of data experimentation is going to be heavily dependent on privacy preserving techniques mainly as it{'}s going to…

Privacy Preserving

Privacy-Preserving Distributed Clustering for Electrical Load Profiling

2020-02-26

Electrical load profiling supports retailers and distribution network operators in having a better understanding of the consumption behavior of consumers. However, traditional clustering methods for load profiling are ce…

ClusteringPrivacy Preserving