Assigning roles to protein mentions: the case of transcription factors.

Transcription factors (TFs) play a crucial role in gene regulation, and providing structured and curated information about them is important for genome biology. Manual curation of TF related data is time-consuming and always lags behind the actual knowledge available in the biomedical literature. He...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of Biomedical Informatics Vol. 42; no. 5; pp. 887 - 895
Autores principales: Yang H, Keane J, Bergman CM, Nenadic G, Yang, Hui, Keane, John, Bergman, Casey M, Nenadic, Goran
Formato: research Journal Article
Publicado: Academic Press Inc. Oct2009
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=105228749&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 105228749
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        15320464
        OMB
      jtl: Journal of Biomedical Informatics
      issn: 15320464
      maglogo: N
    pubinfo:
      dt: Oct2009
      vid: 42
      iid: 5
      pid: 735
      pub: Academic Press Inc.
      place: Burlington, Massachusetts
    artinfo:
      ui:
        105228749
        NLM19364541
        2010429117
        10.1016/j.jbi.2009.04.001
        NLM19364541
        105228749
      ppf: 887
      ppct: 8
      formats:
      tig:
        atl: Assigning roles to protein mentions: the case of transcription factors.
      aug:
        au:
          Yang H
          Keane J
          Bergman CM
          Nenadic G
          Yang, Hui
          Keane, John
          Bergman, Casey M
          Nenadic, Goran
        affil: School of Computer Science, University of Manchester, UK
      sug:
        subj:
          Artificial Intelligence
          Information Retrieval Methods
          Natural Language Processing
          Information Science Methods
          Proteins
          Models, Statistical
          Newsletters
          Reproducibility of Results
      ab: Transcription factors (TFs) play a crucial role in gene regulation, and providing structured and curated information about them is important for genome biology. Manual curation of TF related data is time-consuming and always lags behind the actual knowledge available in the biomedical literature. Here we present a machine-learning text mining approach for identification and tagging of protein mentions that play a TF role in a given context to support the curation process. More precisely, the method explicitly identifies those protein mentions in text that refer to their potential TF functions. The prediction features are engineered from the results of shallow parsing and domain-specific processing (recognition of relevant appearing in phrases) and a phrase-based Conditional Random Fields (CRF) model is used to capture the content and context information of candidate entities. The proposed approach for the identification of TF mentions has been tested on a set of evidence sentences from the TRANSFAC and FlyTF databases. It achieved an F-measure of around 51.5% with a precision of 62.5% using 5-fold cross-validation evaluation. The experimental results suggest that the phrase-based CRF model benefits from the flexibility to use correlated domain-specific features that describe the dependencies between TFs and other entities. To the best of our knowledge, this work is one of the first attempts to apply text-mining techniques to the task of assigning semantic roles to protein mentions.
      pubtype: Academic Journal
      doctype:
        research
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N