paper-with-me

홈 › Papers

Disclose Models, Hide the Data - How to Make Use of Confidential Corpora without Seeing Sensitive Raw Data

2014-05-01 · LREC 2014 5 · Erik Faessler, Johannes Hellrich, Udo Hahn

Confidential corpora from the medical, enterprise, security or intelligence domains often contain sensitive raw data which lead to severe restrictions as far as the public accessibility and distribution of such language resources are concerned. The enforcement of strict mechanisms of data protection consitutes a serious barrier for progress in language technology (products) in such domains, since these data are extremely rare or even unavailable for scientists and developers not directly involved in the creation and maintenance of such resources. In order to by-pass this problem, we here propose to distribute trained language models which were derived from such resources as a substitute for the original confidential raw data which remain hidden to the outside world. As an example, we exploit the access-protected German-language medical FRAMED corpus from which we generate and distribute models for sentence splitting, tokenization and POS tagging based on software taken from OPENNLP, NLTK and JCORE, our own UIMA-based text analytics pipeline.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

POSPOS TaggingSentence

Similar Papers 제목 키워드 기반

Conflict-free joint sampling for preference satisfaction through quantum interference

2022-08-05 · Hiroaki Shinkawa, Nicolas Chauvet, André Röhm, Takatomo Mihana 외

Collective decision-making is vital for recent information and communications technologies. In our previous research, we mathematically derived conflict-free joint decision-making that optimally satisfies players' probab…

Decision Making

A Comparative Study of Image Disguising Methods for Confidential Outsourced Learning

2022-12-31 · Sagar Sharma, Yuechun Gu, Keke Chen

Large training data and expensive model tweaking are standard features of deep learning for images. As a result, data owners often utilize cloud resources to develop large-scale complex models, which raises privacy conce…

Deep Learning feature selection to unhide demographic recommender systems factors

2020-06-17 · Jesús Bobadilla, Ángel González-Prieto, Fernando Ortega, Raúl Lara-Cabrera

Extracting demographic features from hidden factors is an innovative concept that provides multiple and relevant applications. The matrix factorization model generates factors which do not incorporate semantic knowledge.…

Collaborative FilteringDeep LearningFairnessfeature selection+2

Proteus: Preserving Model Confidentiality during Graph Optimizations

2024-04-18 · Yubo Gao, Maryam Haghifam, Christina Giannoula, Renbo Tu 외

Deep learning (DL) models have revolutionized numerous domains, yet optimizing them for computational efficiency remains a challenging endeavor. Development of new DL models typically involves two parties: the model deve…

Computational EfficiencymodelModel Optimization

Data privacy protection in microscopic image analysis for material data mining

2021-11-09 · Boyuan Ma, Xiang Yin, Xiaojuan Ban, Haiyou Huang 외

Recent progress in material data mining has been driven by high-capacity models trained on large datasets. However, collecting experimental data has been extremely costly owing to the amount of human effort and expertise…

Federated LearningImage SegmentationSemantic SegmentationStyle Transfer