Improving rare disease classification using imperfect knowledge graph.

Background: Accurately recognizing rare diseases based on symptom description is an important task in patient triage, early risk stratification, and target therapies. However, due to the very nature of rare diseases, the lack of historical data poses a great challenge to machine learning-based appro...

Descripción completa

Detalles Bibliográficos
Publicado en:BMC Medical Informatics & Decision Making Vol. 19
Autores principales: Li, Xuedong, Wang, Yue, Wang, Dongwu, Yuan, Walter, Peng, Dezhong, Mei, Qiaozhu
Formato: research Journal Article
Publicado: BioMed Central 12/5/2019 Supplement 5
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=140156071&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 140156071
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        14726947
        1CI0
      jtl: BMC Medical Informatics & Decision Making
      issn: 14726947
      maglogo: N
    pubinfo:
      dt: 12/5/2019 Supplement 5
      vid: 19
      pid: 24147
      pub: BioMed Central
    artinfo:
      ui:
        140156071
        140156071
        NLM31801534
        140156071
        10.1186/s12911-019-0938-1
        NLM31801534
        140156071
      ppct: 1
      formats:
      tig:
        atl: Improving rare disease classification using imperfect knowledge graph.
      aug:
        au:
          Li, Xuedong
          Wang, Yue
          Wang, Dongwu
          Yuan, Walter
          Peng, Dezhong
          Mei, Qiaozhu
        affil: College of Computer Science, Sichuan University, Chengdu, China
      sug:
        subj:
          Disease Attributes Classification
          Algorithms
          Human
          Information Science
          Triage
          Validation Studies
          Comparative Studies
          Evaluation Research
          Multicenter Studies
      ab: Background: Accurately recognizing rare diseases based on symptom description is an important task in patient triage, early risk stratification, and target therapies. However, due to the very nature of rare diseases, the lack of historical data poses a great challenge to machine learning-based approaches. On the other hand, medical knowledge in automatically constructed knowledge graphs (KGs) has the potential to compensate the lack of labeled training examples. This work aims to develop a rare disease classification algorithm that makes effective use of a knowledge graph, even when the graph is imperfect.Method: We develop a text classification algorithm that represents a document as a combination of a "bag of words" and a "bag of knowledge terms," where a "knowledge term" is a term shared between the document and the subgraph of KG relevant to the disease classification task. We use two Chinese disease diagnosis corpora to evaluate the algorithm. The first one, HaoDaiFu, contains 51,374 chief complaints categorized into 805 diseases. The second data set, ChinaRe, contains 86,663 patient descriptions categorized into 44 disease categories.Results: On the two evaluation data sets, the proposed algorithm delivers robust performance and outperforms a wide range of baselines, including resampling, deep learning, and feature selection approaches. Both classification-based metric (macro-averaged F1 score) and ranking-based metric (mean reciprocal rank) are used in evaluation.Conclusion: Medical knowledge in large-scale knowledge graphs can be effectively leveraged to improve rare diseases classification models, even when the knowledge graph is incomplete.
      pubtype: Academic Journal
      doctype:
        research
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N