Describir: Language chunking, data sparseness, and the value of a long marker list: explorations with word n-grams and authorial attribution.