Identification and classification of DICOM files with burned-in text content.

Background: Protected health information burned in pixel data is not indicated for various reasons in DICOM. It complicates the secondary use of such data. In recent years, there have been several attempts to anonymize or de-identify DICOM files. Existing approaches have different constraints. No co...

Descripción completa

Detalles Bibliográficos
Publicado en:International Journal of Medical Informatics Vol. 126; pp. 128 - 138
Autores principales: Vcelak, Petr, Kryl, Martin, Kratochvil, Michal, Kleckova, Jana
Formato: research Journal Article
Publicado: Elsevier B.V. Jun2019
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=136017570&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 136017570
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        13865056
        JR4
      jtl: International Journal of Medical Informatics
      issn: 13865056
      maglogo: N
    pubinfo:
      dt: Jun2019
      vid: 126
      pid: 467
      pub: Elsevier B.V.
      place: New York, New York
    artinfo:
      ui:
        136017570
        136017570
        NLM31029254
        136017570
        10.1016/j.ijmedinf.2019.02.011
        NLM31029254
        136017570
      ppf: 128
      ppct: 10
      formats:
      tig:
        atl: Identification and classification of DICOM files with burned-in text content.
      aug:
        au:
          Vcelak, Petr
          Kryl, Martin
          Kratochvil, Michal
          Kleckova, Jana
        affil: NTIS – New Technologies for the Information Society, University of West Bohemia, Univerzitni 8, 30614 Plzen, Czech Republic
      sug:
        subj:
          Privacy and Confidentiality
          Data Security
          Algorithms
          United States
          Human
          Data Collection
          Health Insurance Portability and Accountability Act
          Validation Studies
          Comparative Studies
          Evaluation Research
          Multicenter Studies
      ab: Background: Protected health information burned in pixel data is not indicated for various reasons in DICOM. It complicates the secondary use of such data. In recent years, there have been several attempts to anonymize or de-identify DICOM files. Existing approaches have different constraints. No completely reliable solution exists. Especially for large datasets, it is necessary to quickly analyse and identify files potentially violating privacy.Methods: Classification is based on adaptive-iterative algorithm designed to identify one of three classes. There are several image transformations, optical character recognition, and filters; then a local decision is made. A confirmed local decision is the final one. The classifier was trained on a dataset composed of 15,334 images of various modalities.Results: The false positive rates are in all cases below 4.00%, and 1.81% in the mission-critical problem of detecting protected health information. The classifier's weighted average recall was 94.85%, the weighted average inverse recall was 97.42% and Cohen's Kappa coefficient was 0.920.Conclusion: The proposed novel approach for classification of burned-in text is highly configurable and able to analyse images from different modalities with a noisy background. The solution was validated and is intended to identify DICOM files that need to have restricted access or be thoroughly de-identified due to privacy issues. Unlike with existing tools, the recognised text, including its coordinates, can be further used for de-identification.
      pubtype: Academic Journal
      doctype:
        research
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N