Describir: Efficient feature selection and classification of protein sequence data in bioinformatics.