De-identification In practice
We report our effort to identify the sensitive information, subset of data items listed by HIPAA (Health Insurance Portability and Accountability), from medical text using the recent advances in natural language processing and machine learning techniques. We represent the words with high dimensional continuous vectors learned by a variant of Word2Vec called Continous Bag Of Words (CBOW). We feed the word vectors into a simple neural network with a Long Short-Term Memory (LSTM) architecture. Without any attempts to extract manually crafted features and considering that our medical dataset is too small to be fed into neural network, we obtained promising results. The results thrilled us to think about the larger scale of the project with precise parameter tuning and other possible improvements.
Code (0)
등록된 구현이 없습니다.
Tasks
De-identificationSimilar Papers 제목 키워드 기반
Towards an efficient and risk aware strategy for guiding farmers in identifying best crop management
Identification of best performing fertilizer practices among a set of contrasting practices with field trials is challenging as crop losses are costly for farmers. To identify best management practices, an ''intuitive st…
ManagementRe-ID done right: towards good practices for person re-identification
Training a deep architecture using a ranking loss has become standard for the person re-identification task. Increasingly, these deep architectures include additional components that leverage part detections, attribute p…
AttributePerson Re-IdentificationOnline Multi-Object Tracking with Unsupervised Re-Identification Learning and Occlusion Estimation
Occlusion between different objects is a typical challenge in Multi-Object Tracking (MOT), which often leads to inferior tracking results due to the missing detected objects. The common practice in multi-object tracking …
Multi-Object TrackingObjectObject TrackingOcclusion Estimation+1Computer-Suported Risk Identification for the Holistic Management of Risks
Risk is part of the fabric of every business; surprisingly, there is little work on establishing best practices for systematic, repeatable risk identification, arguably the first step of any risk management process. In t…
ManagementLegal Framework, Dataset and Annotation Schema for Socially Unacceptable Online Discourse Practices in Slovene
In this paper we present the legal framework, dataset and annotation schema of socially unacceptable discourse practices on social networking platforms in Slovenia. On this basis we aim to train an automatic identificati…
General Classification