Unsupervised identification of text reuse in early Chinese literature.

Text reuse in early Chinese transmitted texts is extensive and widespread, often reflecting complex textual histories involving repeated transcription, compilation, and editing spanning many centuries and involving the work of multiple authors and editors. In this study, a fully automated method of...

Descripción completa

Detalles Bibliográficos
Publicado en:Digital Scholarship in the Humanities Vol. 33; no. 3; pp. 670 - 685
Autor principal: Sturgeon, Donald
Formato: Artículo
Publicado: Oxford University Press / USA Sep2018
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=131417032&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 131417032
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        2055768X
        JEO9
      jtl: Digital Scholarship in the Humanities
      issn: 2055768X
      maglogo: N
    pubinfo:
      dt: Sep2018
      vid: 33
      iid: 3
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        131417032
        10.1093/llc/fqx024
      ppf: 670
      ppct: 15
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 969KB
      tig:
        atl: Unsupervised identification of text reuse in early Chinese literature.
      aug:
        au: Sturgeon, Donald
        affil: Fairbank Center for Chinese Studies, Harvard University, USA
      su:
        Chinese literature
        Online databases
        Data mining
        Corpora
        Stemming (Linguistics)
      sug:
        subj:
          Chinese literature
          Online databases
          Data mining
          Corpora
          Stemming (Linguistics)
      ab: Text reuse in early Chinese transmitted texts is extensive and widespread, often reflecting complex textual histories involving repeated transcription, compilation, and editing spanning many centuries and involving the work of multiple authors and editors. In this study, a fully automated method of identifying and representing complex text reuse patterns is presented, and the results evaluated by comparison to a manually compiled reference work. The resultant data are integrated into a widely used and publicly available online database system with browse, search, and visualization functionality. These same results are then aggregated to create a model of text reuse relationships at a corpus level, revealing patterns of systematic reuse among groups of texts. Lastly, the large number of reuse instances identified make possible the analysis of frequently observed string substitutions, which are observed to be strongly indicative of partial synonymy between strings.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Digital Scholarship in the Humanities
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2018
    holdings:
      @attributes:
        islocal: N