Computational predictions for protein sequences of COVID-19 virus via machine learning algorithms.

The rapid spread of coronavirus disease (COVID-19) has become a worldwide pandemic and affected more than 15 million patients reported in 27 countries. Therefore, the computational biology carrying this virus that correlates with the human population urgently needs to be understood. In this paper, t...

Descripción completa

Detalles Bibliográficos
Publicado en:Medical & Biological Engineering & Computing Vol. 59; no. 9; pp. 1723 - 1735
Autores principales: Afify, Heba M., Zanaty, Muhammad S.
Formato: Journal Article
Publicado: Springer Nature Sep2021
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=152044049&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 152044049
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        01400118
        PO0
      jtl: Medical & Biological Engineering & Computing
      issn: 01400118
      maglogo: N
    pubinfo:
      dt: Sep2021
      vid: 59
      iid: 9
      pid: 237
      pub: Springer Nature
      place: New York, New York
    artinfo:
      ui:
        152044049
        151501880
        152044049
        NLM34291385
        10.1007/s11517-021-02412-z
        NLM34291385
        152044049
      ppf: 1723
      ppct: 12
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
      tig:
        atl: Computational predictions for protein sequences of COVID-19 virus via machine learning algorithms.
      aug:
        au:
          Afify, Heba M.
          Zanaty, Muhammad S.
        affil: Systems and Biomedical Engineering Department, Higher Institute of Engineering in El-Shorouk City, Cairo, Egypt
      sug:
      ab: The rapid spread of coronavirus disease (COVID-19) has become a worldwide pandemic and affected more than 15 million patients reported in 27 countries. Therefore, the computational biology carrying this virus that correlates with the human population urgently needs to be understood. In this paper, the classification of the human protein sequences of COVID-19, according to the country, is presented based on machine learning algorithms. The proposed model is based on distinguishing 9238 sequences using three stages, including data preprocessing, data labeling, and classification. In the first stage, data preprocessing's function converts the amino acids of COVID-19 protein sequences into eight groups of numbers based on the amino acids' volume and dipole. It is based on the conjoint triad (CT) method. In the second stage, there are two methods for labeling data from 27 countries from 0 to 26. The first method is based on selecting one number for each country according to the code numbers of countries, while the second method is based on binary elements for each country. According to their countries, machine learning algorithms are used to discover different COVID-19 protein sequences in the last stage. The obtained results demonstrate 100% accuracy, 100% sensitivity, and 90% specificity via the country-based binary labeling method with a linear support vector machine (SVM) classifier. Furthermore, with significant infection data, the USA is more prone to correct classification compared to other countries with fewer data. The unbalanced data for COVID-19 protein sequences is considered a major issue, especially as the US's available data represents 76% of a total of 9238 sequences. The proposed model will act as a prediction tool for the COVID-19 protein sequences in different countries.
      pubtype: Academic Journal
      doctype: Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N