A machine learning based approach to identify protected health information in Chinese clinical text.

Background: With the increasing application of electronic health records (EHRs) in the world, protecting private information in clinical text has drawn extensive attention from healthcare providers to researchers. De-identification, the process of identifying and removing protected health informatio...

Descripción completa

Detalles Bibliográficos
Publicado en:International Journal of Medical Informatics Vol. 116; pp. 24 - 33
Autores principales: Du, Liting, Xia, Chenxi, Deng, Zhaohua, Lu, Gary, Xia, Shuxu, Ma, Jingdong
Formato: research Journal Article
Publicado: Elsevier B.V. Aug2018
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=130046173&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 130046173
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        13865056
        JR4
      jtl: International Journal of Medical Informatics
      issn: 13865056
      maglogo: N
    pubinfo:
      dt: Aug2018
      vid: 116
      pid: 467
      pub: Elsevier B.V.
      place: New York, New York
    artinfo:
      ui:
        130046173
        130046173
        NLM29887232
        130046173
        10.1016/j.ijmedinf.2018.05.010
        NLM29887232
        130046173
      ppf: 24
      ppct: 9
      formats:
      tig:
        atl: A machine learning based approach to identify protected health information in Chinese clinical text.
      aug:
        au:
          Du, Liting
          Xia, Chenxi
          Deng, Zhaohua
          Lu, Gary
          Xia, Shuxu
          Ma, Jingdong
        affil: School of Medicine and Health Management, Tongji Medical College, Huazhong University of Science and Technology, Hubei, China
      sug:
        subj:
          Data Security
          Privacy and Confidentiality
          China
          Software
          Algorithms
          Natural Language Processing
          Human
          Validation Studies
          Comparative Studies
          Evaluation Research
          Multicenter Studies
      ab: Background: With the increasing application of electronic health records (EHRs) in the world, protecting private information in clinical text has drawn extensive attention from healthcare providers to researchers. De-identification, the process of identifying and removing protected health information (PHI) from clinical text, has been central to the discourse on medical privacy since 2006. While de-identification is becoming the global norm for handling medical records, there is a paucity of studies on its application on Chinese clinical text. Without efficient and effective privacy protection algorithms in place, the use of indispensable clinical information would be confined.Objectives: We aimed to (i) describe the current process for PHI in China, (ii) propose a machine learning based approach to identify PHI in Chinese clinical text, and (iii) validate the effectiveness of the machine learning algorithm for de-identification in Chinese clinical text.Methods: Based on 14,719 discharge summaries from regional health centers in Ya'an City, Sichuan province, China, we built a conditional random fields (CRF) model to identify PHI in clinical text, and then used the regular expressions to optimize the recognition results of the PHI categories with fewer samples.Results: We constructed a Chinese clinical text corpus with PHI tags through substantial manual annotation, wherein the descriptive statistics of PHI manifested its wide range and diverse categories. The evaluation showed with a high F-measure of 0.9878 that our CRF-based model had a good performance for identifying PHI in Chinese clinical text.Conclusion: The rapid adoption of EHR in the health sector has created an urgent need for tools that can parse patient specific information from Chinese clinical text. Our application of CRF algorithms for de-identification has shown the potential to meet this need by offering a highly accurate and flexible solution to analyzing Chinese clinical text.
      pubtype: Academic Journal
      doctype:
        research
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N