Knowledge discovery and sequence-based prediction of pandemic influenza using an integrated classification and association rule mining (CBA) algorithm.
Pandemic influenza is a major concern worldwide. Availability of advanced technologies and the nucleotide sequences of a large number of pandemic and non-pandemic influenza viruses in 2009 provide a great opportunity to investigate the underlying rules of pandemic induction through data mining tools...
| Publicado en: | Journal of Biomedical Informatics Vol. 57; pp. 181 - 189 |
|---|---|
| Autores principales: | , , |
| Formato: | research Journal Article |
| Publicado: |
Academic Press Inc.
Oct2015
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=110866340&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 110866340 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 15320464 OMB jtl: Journal of Biomedical Informatics issn: 15320464 maglogo: N pubinfo: dt: Oct2015 vid: 57 pid: 735 pub: Academic Press Inc. place: Burlington, Massachusetts artinfo: ui: 110866340 110866340 NLM26232668 110866340 10.1016/j.jbi.2015.07.018 NLM26232668 110866340 ppf: 181 ppct: 8 formats: tig: atl: Knowledge discovery and sequence-based prediction of pandemic influenza using an integrated classification and association rule mining (CBA) algorithm. aug: au: Kargarfard, Fatemeh Sami, Ashkan Ebrahimie, Esmaeil affil: Department of Computer Science and IT, School of Electrical Engineering and Computer Science, Shiraz University, Shiraz, Iran sug: subj: Influenza A Virus, H1N1 Subtype Disease Outbreaks Algorithms Proteins Data Mining Amino Acids Forecasting Influenza, Human Genetic Techniques ab: Pandemic influenza is a major concern worldwide. Availability of advanced technologies and the nucleotide sequences of a large number of pandemic and non-pandemic influenza viruses in 2009 provide a great opportunity to investigate the underlying rules of pandemic induction through data mining tools. Here, for the first time, an integrated classification and association rule mining algorithm (CBA) was used to discover the rules underpinning alteration of non-pandemic sequences to pandemic ones. We hypothesized that the extracted rules can lead to the development of an efficient expert system for prediction of influenza pandemics. To this end, we used a large dataset containing 5373 HA (hemagglutinin) segments of the 2009 H1N1 pandemic and non-pandemic influenza sequences. The analysis was carried out for both nucleotide and protein sequences. We found a number of new rules which potentially present the undiscovered antigenic sites at influenza structure. At the nucleotide level, alteration of thymine (T) at position 260 was the key discriminating feature in distinguishing non-pandemic from pandemic sequences. At the protein level, rules including I233K, M334L were the differentiating features. CBA efficiently classifies pandemic and non-pandemic sequences with high accuracy at both the nucleotide and protein level. Finding hotspots in influenza sequences is a significant finding as they represent the regions with low antibody reactivity. We argue that the virus breaks host immunity response by mutation at these spots. Based on the discovered rules, we developed the software, "Prediction of Pandemic Influenza" for discrimination of pandemic from non-pandemic sequences. This study opens a new vista in discovery of association rules between mutation points during evolution of pandemic influenza. pubtype: Academic Journal doctype: research Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|