An Effective Big Data Supervised Imbalanced Classification Approach for Ortholog Detection in Related Yeast Species.
Orthology detection requires more effective scaling algorithms. In this paper, a set of gene pair features based on similarity measures (alignment scores, sequence length, gene membership to conserved regions, and physicochemical profiles) are combined in a supervised pairwise ortholog detection app...
| Publicado en: | BioMed Research International Vol. 2015; pp. 1 - 13 |
|---|---|
| Autores principales: | , , , , , |
| Formato: | research Journal Article |
| Publicado: |
Wiley-Blackwell
10/29/2015
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=113629972&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 113629972 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 23146133 FT2T jtl: BioMed Research International issn: 23146133 maglogo: N pubinfo: dt: 10/29/2015 vid: 2015 pid: 480 pub: Wiley-Blackwell place: Malden, Massachusetts artinfo: ui: 113629972 113629972 113629972 10.1155/2015/748681 113629972 ppf: 1 ppct: 12 formats: fmt: @attributes: type: P tig: atl: An Effective Big Data Supervised Imbalanced Classification Approach for Ortholog Detection in Related Yeast Species. aug: au: Galpert, Deborah del Río, Sara Herrera, Francisco Ancede-Gallardo, Evys Antunes, Agostinho Agüero-Chapin, Guillermin affil: Departamento de Ciencias de la Computación, Universidad Central “Marta Abreu” de Las Villas (UCLV), 54830 Santa Clara, Cuba sug: subj: Data Analytics Algorithms Yeasts Biomedical Engineering Genome ab: Orthology detection requires more effective scaling algorithms. In this paper, a set of gene pair features based on similarity measures (alignment scores, sequence length, gene membership to conserved regions, and physicochemical profiles) are combined in a supervised pairwise ortholog detection approach to improve effectiveness considering low ortholog ratios in relation to the possible pairwise comparison between two genomes. In this scenario, big data supervised classifiers managing imbalance between ortholog and nonortholog pair classes allow for an effective scaling solution built from two genomes and extended to other genome pairs. The supervised approach was compared with RBH, RSD, and OMA algorithms by using the following yeast genome pairs: Saccharomyces cerevisiae-Kluyveromyces lactis, Saccharomyces cerevisiae-Candida glabrata, and Saccharomyces cerevisiae-Schizosaccharomyces pombe as benchmark datasets. Because of the large amount of imbalanced data, the building and testing of the supervised model were only possible by using big data supervised classifiers managing imbalance. Evaluation metrics taking low ortholog ratios into account were applied. From the effectiveness perspective, MapReduce Random Oversampling combined with Spark SVM outperformed RBH, RSD, and OMA, probably because of the consideration of gene pair features beyond alignment similarities combined with the advances in big data supervised classification. pubtype: Academic Journal doctype: research Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|