Learning structured medical information from social media.

Our goal is to summarise and aggregate information from social media regarding the symptoms of a disease, the drugs used and the treatment effects both positive and negative. To achieve this we first apply a supervised machine learning method to automatically extract medical concepts from natural la...

Full description

Bibliographic Details
Published in:Journal of Biomedical Informatics Vol. 110
Main Authors: Hasan, Abul, Levene, Mark, Weston, David
Format: Journal Article
Published: Academic Press Inc. Oct2020
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=146481897&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 146481897
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        15320464
        OMB
      jtl: Journal of Biomedical Informatics
      issn: 15320464
      maglogo: N
    pubinfo:
      dt: Oct2020
      vid: 110
      pid: 735
      pub: Academic Press Inc.
      place: Burlington, Massachusetts
    artinfo:
      ui:
        146481897
        146481897
        NLM32942027
        10.1016/j.jbi.2020.103568
        NLM32942027
        146481897
      ppct: 1
      formats:
      tig:
        atl: Learning structured medical information from social media.
      aug:
        au:
          Hasan, Abul
          Levene, Mark
          Weston, David
        affil: Department of Computer Science and Information Systems, Birkbeck, University of London, London WC1E 7HX, United Kingdom
      sug:
        subj:
          Social Media
          Algorithms
          Language
          Social Readjustment Rating Scale
      ab: Our goal is to summarise and aggregate information from social media regarding the symptoms of a disease, the drugs used and the treatment effects both positive and negative. To achieve this we first apply a supervised machine learning method to automatically extract medical concepts from natural language text. In an environment such as social media, where new data is continuously streamed, we need a methodology that will allow us to continuously train with the new data. To attain such incremental re-training, a semi-supervised methodology is developed, which is capable of learning new concepts from a small set of labelled data together with the much larger set of unlabelled data. The semi-supervised methodology deploys a conditional random field (CRF) as the base-line training algorithm for extracting medical concepts. The methodology iteratively augments to the training set sentences having high confidence, and adds terms to existing dictionaries to be used as features with the base-line model for further classification. Our empirical results show that the base-line CRF performs strongly across a range of different dictionary and training sizes; when the base-line is built with the full training data the F1 score reaches the range 84%-90%. Moreover, we show that the semi-supervised method produces a mild but significant improvement over the base-line. We also discuss the significance of the potential improvement of the semi-supervised methodology and found that it is significantly more accurate in most cases than the underlying base-line model.
      pubtype: Academic Journal
      doctype: Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N