| Sumario: | • Chemical elements concentration in ancient pottery measured through non-invasive technique. • Provenance of ancient pottery determined through machine learning classification techniques. • Multiple measurements exploited in presence of inhomogeneous material. • One-class classification techniques useful when training data have the same provenance. Pottery classification based on chemical composition characterization through non-invasive analytical techniques is a well-known method typically adopted to solve the problem of pottery provenance attribution. Machine learning approaches have been recently introduced as a tool to develop models for archaeological classification directly inferred from data. This classification is quite often given in terms of local or non-local samples: in these cases, if the hypothesized provenance is to be validated, one-class models could be more adequate than binary or multi-class predictors. Indeed, one-class classifiers are trained only using positive examples of the class to be learned, and they can be subsequently tested on positive and negative examples in order to evaluate their generalization capability. They are thus in principle more apt to efficiently classify local samples, as the non-local ones do not naturally gather in a well-defined class. In this paper, we tested a one-class classifier on a dataset of 112 examples representing pottery fragments described in terms of nine chemical elements. Different examples of the dataset can correspond to measurements done on a same physical fragment, thus the hypothesis of independence among observations might be violated. For this reason, we investigated the use of three data stratification techniques, based on physical fragments and on measures. We employed the support vector one-class classification algorithm, finding the smallest sphere in a feature space that contains most of the training points. The obtained classification performances were compared with those of several machine learning algorithms for binary classification. All the models were trained by a nested cross validation technique, separately taking into account the fine-tuning of hyperparameters and the robust estimation of generalization performance. Comparisons were done on the same dataset according to different performance metrics. The obtained results show that one-class classification attains a similar sensitivity of binary classification approaches, meanwhile improving the performance in terms of specificity, therefore showing a good behavior both on positive and negative examples.
|