Identification and classification of DICOM files with burned-in text content.
Background: Protected health information burned in pixel data is not indicated for various reasons in DICOM. It complicates the secondary use of such data. In recent years, there have been several attempts to anonymize or de-identify DICOM files. Existing approaches have different constraints. No co...
| Publicado en: | International Journal of Medical Informatics Vol. 126; pp. 128 - 138 |
|---|---|
| Autores principales: | , , , |
| Formato: | research Journal Article |
| Publicado: |
Elsevier B.V.
Jun2019
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=136017570&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 136017570 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 13865056 JR4 jtl: International Journal of Medical Informatics issn: 13865056 maglogo: N pubinfo: dt: Jun2019 vid: 126 pid: 467 pub: Elsevier B.V. place: New York, New York artinfo: ui: 136017570 136017570 NLM31029254 136017570 10.1016/j.ijmedinf.2019.02.011 NLM31029254 136017570 ppf: 128 ppct: 10 formats: tig: atl: Identification and classification of DICOM files with burned-in text content. aug: au: Vcelak, Petr Kryl, Martin Kratochvil, Michal Kleckova, Jana affil: NTIS – New Technologies for the Information Society, University of West Bohemia, Univerzitni 8, 30614 Plzen, Czech Republic sug: subj: Privacy and Confidentiality Data Security Algorithms United States Human Data Collection Health Insurance Portability and Accountability Act Validation Studies Comparative Studies Evaluation Research Multicenter Studies ab: Background: Protected health information burned in pixel data is not indicated for various reasons in DICOM. It complicates the secondary use of such data. In recent years, there have been several attempts to anonymize or de-identify DICOM files. Existing approaches have different constraints. No completely reliable solution exists. Especially for large datasets, it is necessary to quickly analyse and identify files potentially violating privacy.Methods: Classification is based on adaptive-iterative algorithm designed to identify one of three classes. There are several image transformations, optical character recognition, and filters; then a local decision is made. A confirmed local decision is the final one. The classifier was trained on a dataset composed of 15,334 images of various modalities.Results: The false positive rates are in all cases below 4.00%, and 1.81% in the mission-critical problem of detecting protected health information. The classifier's weighted average recall was 94.85%, the weighted average inverse recall was 97.42% and Cohen's Kappa coefficient was 0.920.Conclusion: The proposed novel approach for classification of burned-in text is highly configurable and able to analyse images from different modalities with a noisy background. The solution was validated and is intended to identify DICOM files that need to have restricted access or be thoroughly de-identified due to privacy issues. Unlike with existing tools, the recognised text, including its coordinates, can be further used for de-identification. pubtype: Academic Journal doctype: research Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|