Building a specialized lexicon for breast cancer clinical trial subject eligibility analysis.
A natural language processing (NLP) application requires sophisticated lexical resources to support its processing goals. Different solutions, such as dictionary lookup and MetaMap, have been proposed in the healthcare informatics literature to identify disease terms with more than one word (multi-g...
| Publicado en: | Health Informatics Journal Vol. 27; no. 1; pp. 1 - 16 |
|---|---|
| Autores principales: | , , , |
| Formato: | research tables/charts Journal Article |
| Publicado: |
Sage Publications Inc.
Jan-Mar2021
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=150395752&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 150395752 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 14604582 EJK jtl: Health Informatics Journal issn: 14604582 maglogo: Y pubinfo: dt: Jan-Mar2021 vid: 27 iid: 1 pid: 344 pub: Sage Publications Inc. place: Thousand Oaks, California artinfo: ui: 150395752 150395752 150395752 10.1177/1460458221989392 150395752 ppf: 1 ppct: 15 formats: tig: atl: Building a specialized lexicon for breast cancer clinical trial subject eligibility analysis. aug: au: Jung, Euisung Jain, Hemant Sinha, Atish P Gaudioso, Carmelo affil: Information Operations and Technology Management, John B. and Lillian E. Neff College of Business and Innovation, The University of Toledo, USA. sug: subj: Breast Neoplasms Clinical Trials Research Subject Recruitment Eligibility Determination Dictionaries Vocabulary, Controlled Natural Language Processing Data Mining Database Construction Human Information Retrieval Information Resources Online Services National Cancer Institute (U.S.) American Cancer Society World Wide Web Nomenclature Electronic Data Interchange Health Informatics ab: A natural language processing (NLP) application requires sophisticated lexical resources to support its processing goals. Different solutions, such as dictionary lookup and MetaMap, have been proposed in the healthcare informatics literature to identify disease terms with more than one word (multi-gram disease named entities). Although a lot of work has been done in the identification of protein- and gene-named entities in the biomedical field, not much research has been done on the recognition and resolution of terminologies in the clinical trial subject eligibility analysis. In this study, we develop a specialized lexicon for improving NLP and text mining analysis in the breast cancer domain, and evaluate it by comparing it with the Systematized Nomenclature of Medicine Clinical Terms (SNOMED CT). We use a hybrid methodology, which combines the knowledge of domain experts, terms from multiple online dictionaries, and the mining of text from sample clinical trials. Use of our methodology introduces 4243 unique lexicon items, which increase bigram entity match by 38.6% and trigram entity match by 41%. Our lexicon, which adds a significant number of new terms, is very useful for matching patients to clinical trials automatically based on eligibility matching. Beyond clinical trial matching, the specialized lexicon developed in this study could serve as a foundation for future healthcare text mining applications. pubtype: Academic Journal doctype: research tables/charts Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|