paper-with-me

홈 › Papers

Automated Anonymization as Spelling Variant Detection

2016-12-01 · WS 2016 12 · Steven Kester Yuwono, Hwee Tou Ng, Kee Yuan Ngiam

The issue of privacy has always been a concern when clinical texts are used for research purposes. Personal health information (PHI) (such as name and identification number) needs to be removed so that patients cannot be identified. Manual anonymization is not feasible due to the large number of clinical texts to be anonymized. In this paper, we tackle the task of anonymizing clinical texts written in sentence fragments and which frequently contain symbols, abbreviations, and misspelled words. Our clinical texts therefore differ from those in the i2b2 shared tasks which are in prose form with complete sentences. Our clinical texts are also part of a structured database which contains patient name and identification number in structured fields. As such, we formulate our anonymization task as spelling variant detection, exploiting patients{'} personal information in the structured fields to detect their spelling variants in clinical texts. We successfully anonymized clinical texts consisting of more than 200 million words, using minimum edit distance and regular expression patterns.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Sentence

Similar Papers 제목 키워드 기반

Detecting spelling variants in non-standard texts

2017-04-01 · EACL 2017 4 · Fabian Barteld

Spelling variation in non-standard language, e.g. computer-mediated communication and historical texts, is usually treated as a deviation from a standard spelling, e.g. 2mr as an non-standard spelling for tomorrow. Conse…

Context-Sensitive Malicious Spelling Error Correction

2019-01-23 · Hongyu Gong, Yuchen Li, Suma Bhat, Pramod Viswanath

Misspelled words of the malicious kind work by changing specific keywords and are intended to thwart existing automated applications for cyber-environment control such as harassing content detection on the Internet and e…

Spam detectionSpelling CorrectionWord Embeddings

Automated Spelling Correction for Clinical Text Mining in Russian

2020-04-10 · Ksenia Balabaeva, Anastasia Funkner, Sergey Kovalchuk

The main goal of this paper is to develop a spell checker module for clinical text in Russian. The described approach combines string distance measure algorithms with technics of machine learning embedding methods. Our o…

BIG-bench Machine LearningNegationSpelling Correction

No Intruder, no Validity: Evaluation Criteria for Privacy-Preserving Text Anonymization

2021-03-16 · Maximilian Mozes, Bennett Kleinberg

For sensitive text data to be shared among NLP researchers and practitioners, shared documents need to comply with data protection and privacy laws. There is hence a growing interest in automated approaches for text anon…

AttributePrivacy PreservingText Anonymization

Cysill Ar-lein: A Corpus of Written Contemporary Welsh Compiled from an On-line Spelling and Grammar Checker

2016-05-01 · LREC 2016 5 · Delyth Prys, Gruffudd Prys, Dewi Bryn Jones

This paper describes the use of a free, on-line language spelling and grammar checking aid as a vehicle for the collection of a significant (31 million words and rising) corpus of text for academic research in the contex…

Language Modelling