Words in the hood: a new look at the distribution of words in texts.
The problem of repeated occurrences and of the proximity between lexemes in text traditionally has been approached in order to account for certain close distributions such as bursts (rafales) or block-effects. Here we present the initial stages and results of research addressing the same question fr...
| Publicado en: | Literary & Linguistic Computing Vol. 12; no. 2; pp. 71 - 79 |
|---|---|
| Autores principales: | , |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
1997
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=80117252&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 80117252 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 02681145 BJ1 jtl: Literary & Linguistic Computing issn: 02681145 maglogo: N pubinfo: dt: 1997 vid: 12 iid: 2 pid: 622 pub: Oxford University Press / USA artinfo: ui: 80117252 10.1093/llc/12.2.71 ppf: 71 ppct: 8 formats: fmt: @attributes: type: P size: 668KB tig: atl: Words in the hood: a new look at the distribution of words in texts. aug: au: Juillard, M Luong, X affil: CNRS, Universite Nice-Sophia Antipolis, UER Lettres, 98 boulevard E. Herriot, B.P. 209, 06204 Nice cedex 3, France Corresponding author E-mail: juillard@unice.fr su: Topology Multidimensional databases Information literacy Lexical access Grammar sug: subj: Topology Multidimensional databases Information literacy Lexical access Grammar ab: The problem of repeated occurrences and of the proximity between lexemes in text traditionally has been approached in order to account for certain close distributions such as bursts (rafales) or block-effects. Here we present the initial stages and results of research addressing the same question from a totally different angle and with new conceptual tools. The two chief notions of topology, neighbourhood and equivalence of shape, are used to explore systematically the vicinity of lexemes after defining a neighbourhood base. The distribution of a given lexeme is thus associated with a sequence of addresses corresponding to its neighbourhoods. This sequence can be transformed into a characteristic vector V carrying valuable information that can be processed by a number of methods, particularly tree-analysis and multidimensional scaling. We illustrate, with various examples, the successive stages of the method applied to the tagged version of the LOB Corpus of British texts. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Literary & Linguistic Computing holder: Oxford University Press / USA dt: @attributes: year: 1997 holdings: @attributes: islocal: N |
|---|