| Sumario: | The problem of repeated occurrences and of the proximity between lexemes in text traditionally has been approached in order to account for certain close distributions such as bursts (rafales) or block-effects. Here we present the initial stages and results of research addressing the same question from a totally different angle and with new conceptual tools. The two chief notions of topology, neighbourhood and equivalence of shape, are used to explore systematically the vicinity of lexemes after defining a neighbourhood base. The distribution of a given lexeme is thus associated with a sequence of addresses corresponding to its neighbourhoods. This sequence can be transformed into a characteristic vector V carrying valuable information that can be processed by a number of methods, particularly tree-analysis and multidimensional scaling. We illustrate, with various examples, the successive stages of the method applied to the tagged version of the LOB Corpus of British texts.
|