DataSifterText: Partially Synthetic Text Generation for Sensitive Clinical Notes.
Petabytes of health data are collected annually across the globe in electronic health records (EHR), including significant information stored as unstructured free text. However, the lack of effective mechanisms to securely share clinical text has inhibited its full utilization. We propose a new meth...
| Published in: | Journal of Medical Systems Vol. 46; no. 12; pp. 1 - 15 |
|---|---|
| Main Authors: | , , , , |
| Format: | equations & formulas research tables/charts Journal Article |
| Published: |
Springer Nature
Dec2022
|
| Online Access: | View this record in EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=160563226&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 160563226 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 01485598 4N0 jtl: Journal of Medical Systems issn: 01485598 maglogo: N pubinfo: dt: Dec2022 vid: 46 iid: 12 pid: 237 pub: Springer Nature place: New York, New York artinfo: ui: 160563226 160563226 160563226 10.1007/s10916-022-01880-6 160563226 ppf: 1 ppct: 14 formats: fmt: – @attributes: type: T – @attributes: type: P tig: atl: DataSifterText: Partially Synthetic Text Generation for Sensitive Clinical Notes. aug: au: Zhou, Nina Wu, Qiucheng Wu, Zewen Marino, Simeone Dinov, Ivo D. affil: Statistics Online Computational Resource, Health Behavior and Biological, and Department of Biostatistics, University of Michigan, Ann Arbor, USA sug: subj: Data Management Methods Databases, Health Electronic Data Interchange Data Mining Data Collection, Computer Assisted Human Data Science Artificial Intelligence Machine Learning Algorithms Sample Size Data Security Privacy and Confidentiality Probability Quality Assessment Data Analytics Descriptive Statistics Funding Source ab: Petabytes of health data are collected annually across the globe in electronic health records (EHR), including significant information stored as unstructured free text. However, the lack of effective mechanisms to securely share clinical text has inhibited its full utilization. We propose a new method, DataSifterText, to generate partially synthetic clinical free-text that can be safely shared between stakeholders (e.g., clinicians, STEM researchers, engineers, analysts, and healthcare providers), limiting the re-identification risk while providing significantly better utility preservation than suppressing or generalizing sensitive tokens. The method creates partially synthetic free-text data, which inherits the joint population distribution of the original data, and disguises the location of true and obfuscated words. Under certain obfuscation levels, the resulting synthetic text was sufficiently altered with different choices, orders, and frequencies of words compared to the original records. The differences were comparable to machine-generated (fully synthetic) text reported in previous studies. We applied DataSifterText to two medical case studies. In the CDC work injury application, using privacy protection, 60.9-86.5% of the synthetic descriptions belong to the same cluster as the original descriptions, demonstrating better utility preservation than the naïve content suppressing method (45.8-85.7%). In the MIMIC III application, the generated synthetic data maintained over 80% of the original information regarding patients' overall health conditions. The reported DataSifterText statistical obfuscation results indicate that the technique provides sufficient privacy protection (low identification risk) while preserving population-level information (high utility). pubtype: Academic Journal doctype: equations & formulas research tables/charts Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|