Document dissimilarity within and across languages: A benchmarking study.

Quantifying the similarity or dissimilarity between documents is an important task in authorship attribution, information retrieval, plagiarism detection, text mining, and many other areas of linguistic computing. Numerous similarity indices have been devised and used, but relatively little attentio...

Descripción completa

Detalles Bibliográficos
Publicado en:Literary & Linguistic Computing Vol. 29; no. 1; pp. 6 - 23
Autores principales: Forsyth, Richard S., Sharoff, Serge
Formato: Artículo
Publicado: Oxford University Press / USA Apr2014
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=95114212&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 95114212
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        02681145
        BJ1
      jtl: Literary & Linguistic Computing
      issn: 02681145
      maglogo: N
    pubinfo:
      dt: Apr2014
      vid: 29
      iid: 1
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        95114212
        10.1093/llc/fqt002
      ppf: 6
      ppct: 17
      formats:
        fmt:
          @attributes:
            type: P
            size: 736KB
      tig:
        atl: Document dissimilarity within and across languages: A benchmarking study.
      aug:
        au:
          Forsyth, Richard S.
          Sharoff, Serge
        affil: University of Leeds, UK
      su:
        Corpora
        Text processing (Computer science)
        Text mining
        Plagiarism prevention
      sug:
        subj:
          Corpora
          Text processing (Computer science)
          Text mining
          Plagiarism prevention
      ab: Quantifying the similarity or dissimilarity between documents is an important task in authorship attribution, information retrieval, plagiarism detection, text mining, and many other areas of linguistic computing. Numerous similarity indices have been devised and used, but relatively little attention has been paid to calibrating such indices against externally imposed standards, mainly because of the difficulty of establishing agreed reference levels of inter-text similarity. The present article introduces a multi-register corpus gathered for this purpose, in which each text has been located in a similarity space based on ratings by human readers. This provides a resource for testing similarity measures derived from computational text-processing against reference levels derived from human judgement, i.e. external to the texts themselves. We describe the results of a benchmarking study in five different languages in which some widely used measures perform comparatively poorly. In particular, several alternative correlational measures (Pearson r, Spearman rho, tetrachoric correlation) consistently outperform cosine similarity on our data. A method of using what we call ‘anchor texts’ to extend this method from monolingual inter-text similarity-scoring to inter-text similarity-scoring across languages is also proposed and tested.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Literary & Linguistic Computing
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2014
    holdings:
      @attributes:
        islocal: N