FEATURE EXTRACTION MODELS FOR MEDICAL KNOWLEDGE REPRESENTATION AS ENTITY RECOGNITION TO FACT CONSTRUCTION.
As the mass of medical data and its growing availability continue to rise, the difficulty of deriving knowledge out of texts in natural language is getting bigger. To cope with this complexity, the information extraction has become one of the cornerstones of artificial intelligence and text analysis...
| Publicado en: | Scientific Culture Vol. 12; no. 1, Part 1; pp. 2451 - 2467 |
|---|---|
| Autores principales: | , , |
| Formato: | Artículo |
| Publicado: |
University of the Aegean
2026
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=192213588&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 192213588 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 24080071 I6HU jtl: Scientific Culture issn: 24080071 maglogo: N pubinfo: dt: 2026 vid: 12 iid: 1, Part 1 pid: 47715 pub: University of the Aegean artinfo: ui: 192213588 10.5281/zenodo.121126172 ppf: 2451 ppct: 16 formats: tig: atl: FEATURE EXTRACTION MODELS FOR MEDICAL KNOWLEDGE REPRESENTATION AS ENTITY RECOGNITION TO FACT CONSTRUCTION. aug: au: Abdullah, Sura Mahmood Al-Bakry, Abbas Mohsin Farhan, Alaa K. affil: Iraqi Commission for Computers and Informatics/ University of Information Technology and Communication Iraq–Baghdad. University of Information Technology and Communication (UoITC) Iraq-Baghdad. College of Computer Sciences/University of Technology – Iraq- Baghdad. su: Feature extraction Natural language processing Language models Coronary artery disease Machine learning Medical records Data mining sug: subj: Feature extraction Natural language processing Language models Coronary artery disease Machine learning Medical records Data mining keyword: BERT Conditional Random Field Fact Construction N-gram Named Entity Recognition Natural Language Processing TF-IDF ab: As the mass of medical data and its growing availability continue to rise, the difficulty of deriving knowledge out of texts in natural language is getting bigger. To cope with this complexity, the information extraction has become one of the cornerstones of artificial intelligence and text analysis, with the so-called named entity recognition (NER) technology. The aim of the NER is to define the medical concepts and categorize them under predetermined groups that include symptoms, medications, lab tests, and risk factors. NER is regarded as one of the major steps of the Natural Language Processing (NLP) since it assists in analyzing medical texts and formulating the facts, which are commonly represented by a triad (entity, feature, value). In this paper, the author will derive such facts using medical texts by developing three NER models based on three features extraction methods: Rule-based approach, N-grams, TFIDF-ngrams and BERT. The models have had their applied contextual and linguistic analysis to extract the descriptions qualities of each token in the text depending on the type of ex-tractor that is used. These characteristics are subsequently fed to the Enhanced Conditional Random Field (ECRF) classifier, which the token is classified as to be an entity of a specific category. An analysis of the features and values of every entity is obtained based on the surrounding analysis of the token, which enables the fact to be accurately triadic presented. The three models were trained using our data concerning coronary artery disease that was compiled using several sources. Evaluation findings indicated that the proposed BERT-ECRF model performed better than the other models, having 98 extracted entities with the accuracy of 0.986, precision and recall of 0.986, and an F1 score of 0.984. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y dt: @attributes: year: 2026 holdings: @attributes: islocal: N |
|---|