Automated Training for Algorithms That Learn from Genomic Data.

Supervised machine learning algorithms are used by life scientists for a variety of objectives. Expert-curated public gene and protein databases are major resources for gathering data to train these algorithms. While these data resources are continuously updated, generally, these updates are not inc...

Descripción completa

Detalles Bibliográficos
Publicado en:BioMed Research International Vol. 2015; pp. 1 - 10
Autores principales: Cilingir, Gokcen, Broschat, Shira L.
Formato: research tables/charts Journal Article
Publicado: Wiley-Blackwell 1/28/2015
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=109273211&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 109273211
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        23146133
        FT2T
      jtl: BioMed Research International
      issn: 23146133
      maglogo: N
    pubinfo:
      dt: 1/28/2015
      vid: 2015
      pid: 480
      pub: Wiley-Blackwell
      place: Malden, Massachusetts
    artinfo:
      ui:
        109273211
        109273211
        109273211
        10.1155/2015/234236
        109273211
      ppf: 1
      ppct: 9
      formats:
        fmt:
          @attributes:
            type: P
      tig:
        atl: Automated Training for Algorithms That Learn from Genomic Data.
      aug:
        au:
          Cilingir, Gokcen
          Broschat, Shira L.
        affil: School of Electrical Engineering and Computer Science, Washington State University, Pullman, WA 99164, USA
      sug:
        subj:
          Algorithms
          Genomics
          Human
          Research Methodology
      ab: Supervised machine learning algorithms are used by life scientists for a variety of objectives. Expert-curated public gene and protein databases are major resources for gathering data to train these algorithms. While these data resources are continuously updated, generally, these updates are not incorporated into published machine learning algorithms which thereby can become outdated soon after their introduction. In this paper, we propose a new model of operation for supervised machine learning algorithms that learn from genomic data. By defining these algorithms in a pipeline in which the training data gathering procedure and the learning process are automated, one can create a system that generates a classifier or predictor using information available from public resources. The proposed model is explained using three case studies on SignalP, MemLoci, and ApicoAP in which existing machine learning models are utilized in pipelines. Given that the vast majority of the procedures described for gathering training data can easily be automated, it is possible to transform valuable machine learning algorithms into self-evolving learners that benefit from the ever-changing data available for gene products and to develop new machine learning algorithms that are similarly capable.
      pubtype: Academic Journal
      doctype:
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N