Automatic discourse connective detection in biomedical text.
Objective: Relation extraction in biomedical text mining systems has largely focused on identifying clause-level relations, but increasing sophistication demands the recognition of relations at discourse level. A first step in identifying discourse relations involves the detection of discourse conne...
| Publicado en: | Journal of the American Medical Informatics Association Vol. 19; no. 5; pp. 800 - 809 |
|---|---|
| Autores principales: | , , , , , |
| Formato: | research Journal Article |
| Publicado: |
Oxford University Press / USA
Sep2012
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=104360935&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 104360935 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 10675027 FZ9 jtl: Journal of the American Medical Informatics Association issn: 10675027 maglogo: N pubinfo: dt: Sep2012 vid: 19 iid: 5 pid: 622 pub: Oxford University Press / USA artinfo: ui: 104360935 NLM22744958 2011646704 10.1136/amiajnl-2011-000775 NLM22744958 PMC3422833 104360935 ppf: 800 ppct: 9 formats: tig: atl: Automatic discourse connective detection in biomedical text. aug: au: Polepalli Ramesh, Balaji Prasad, Rashmi Miller, Tim Harrington, Brian Yu, Hong Ramesh, Balaji Polepalli affil: Department of Electrical Engineering and Computer Science, University of Wisconsin-Milwaukee, Milwaukee, Wisconsin 53211, USA sug: subj: Data Mining Methods Natural Language Processing Artificial Intelligence Human Algorithms ab: Objective: Relation extraction in biomedical text mining systems has largely focused on identifying clause-level relations, but increasing sophistication demands the recognition of relations at discourse level. A first step in identifying discourse relations involves the detection of discourse connectives: words or phrases used in text to express discourse relations. In this study supervised machine-learning approaches were developed and evaluated for automatically identifying discourse connectives in biomedical text.Materials and Methods: Two supervised machine-learning models (support vector machines and conditional random fields) were explored for identifying discourse connectives in biomedical literature. In-domain supervised machine-learning classifiers were trained on the Biomedical Discourse Relation Bank, an annotated corpus of discourse relations over 24 full-text biomedical articles (~112,000 word tokens), a subset of the GENIA corpus. Novel domain adaptation techniques were also explored to leverage the larger open-domain Penn Discourse Treebank (~1 million word tokens). The models were evaluated using the standard evaluation metrics of precision, recall and F1 scores.Results and Conclusion: Supervised machine-learning approaches can automatically identify discourse connectives in biomedical text, and the novel domain adaptation techniques yielded the best performance: 0.761 F1 score. A demonstration version of the fully implemented classifier BioConn is available at: http://bioconn.askhermes.org. pubtype: Academic Journal doctype: research Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|