| Sumario: | When geographic language variation is examined on a broad level, it is common practice to aggregate large quantities of data to identify dominating structures. This practice, however, runs the risk of neglecting valuable geolinguistic data. This article presents a perspective of dealing with large collections of dialect maps that aims at preserving the distinctiveness of individual maps. Methods derived from spatial statistics and stochastic image analysis are applied to dialect data to suggest a geographically informed procedure for clustering individual dialect maps based on their spatial patterns. This is achieved by automated comparison of statistical properties of the maps. A fuzzy clustering algorithm is employed, which allows the researcher to detect and measure gradual similarities between individual maps rather than form ‘hard’ clusters as in conventional hierarchical clustering. The method can be used for grouping maps based on their spatial similarities around a prototypical spatial pattern, while at the same time allowing for the investigation of semantic/phonetic/ontic, etc. relationships between such spatially related maps. Further, it is argued that these methods offer promising enhancements for quantitative research techniques on geographical language variation by allowing to (automatically and objectively) find patterns in the data and compare external spatial structures to dialect map corpora.
|