FEATURE EXTRACTION MODELS FOR MEDICAL KNOWLEDGE REPRESENTATION AS ENTITY RECOGNITION TO FACT CONSTRUCTION.

As the mass of medical data and its growing availability continue to rise, the difficulty of deriving knowledge out of texts in natural language is getting bigger. To cope with this complexity, the information extraction has become one of the cornerstones of artificial intelligence and text analysis...

Descripción completa

Detalles Bibliográficos
Publicado en:Scientific Culture Vol. 12; no. 1, Part 1; pp. 2451 - 2467
Autores principales: Abdullah, Sura Mahmood, Al-Bakry, Abbas Mohsin, Farhan, Alaa K.
Formato: Artículo
Publicado: University of the Aegean 2026
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=192213588&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 192213588
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        24080071
        I6HU
      jtl: Scientific Culture
      issn: 24080071
      maglogo: N
    pubinfo:
      dt: 2026
      vid: 12
      iid: 1, Part 1
      pid: 47715
      pub: University of the Aegean
    artinfo:
      ui:
        192213588
        10.5281/zenodo.121126172
      ppf: 2451
      ppct: 16
      formats:
      tig:
        atl: FEATURE EXTRACTION MODELS FOR MEDICAL KNOWLEDGE REPRESENTATION AS ENTITY RECOGNITION TO FACT CONSTRUCTION.
      aug:
        au:
          Abdullah, Sura Mahmood
          Al-Bakry, Abbas Mohsin
          Farhan, Alaa K.
        affil:
          Iraqi Commission for Computers and Informatics/ University of Information Technology and Communication Iraq–Baghdad.
          University of Information Technology and Communication (UoITC) Iraq-Baghdad.
          College of Computer Sciences/University of Technology – Iraq- Baghdad.
      su:
        Feature extraction
        Natural language processing
        Language models
        Coronary artery disease
        Machine learning
        Medical records
        Data mining
      sug:
        subj:
          Feature extraction
          Natural language processing
          Language models
          Coronary artery disease
          Machine learning
          Medical records
          Data mining
      keyword:
        BERT
        Conditional Random Field
        Fact Construction
        N-gram
        Named Entity Recognition
        Natural Language Processing
        TF-IDF
      ab: As the mass of medical data and its growing availability continue to rise, the difficulty of deriving knowledge out of texts in natural language is getting bigger. To cope with this complexity, the information extraction has become one of the cornerstones of artificial intelligence and text analysis, with the so-called named entity recognition (NER) technology. The aim of the NER is to define the medical concepts and categorize them under predetermined groups that include symptoms, medications, lab tests, and risk factors. NER is regarded as one of the major steps of the Natural Language Processing (NLP) since it assists in analyzing medical texts and formulating the facts, which are commonly represented by a triad (entity, feature, value). In this paper, the author will derive such facts using medical texts by developing three NER models based on three features extraction methods: Rule-based approach, N-grams, TFIDF-ngrams and BERT. The models have had their applied contextual and linguistic analysis to extract the descriptions qualities of each token in the text depending on the type of ex-tractor that is used. These characteristics are subsequently fed to the Enhanced Conditional Random Field (ECRF) classifier, which the token is classified as to be an entity of a specific category. An analysis of the features and values of every entity is obtained based on the surrounding analysis of the token, which enables the fact to be accurately triadic presented. The three models were trained using our data concerning coronary artery disease that was compiled using several sources. Evaluation findings indicated that the proposed BERT-ECRF model performed better than the other models, having 98 extracted entities with the accuracy of 0.986, precision and recall of 0.986, and an F1 score of 0.984.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      dt:
        @attributes:
          year: 2026
    holdings:
      @attributes:
        islocal: N