Unsupervised identification of text reuse in early Chinese literature.
Text reuse in early Chinese transmitted texts is extensive and widespread, often reflecting complex textual histories involving repeated transcription, compilation, and editing spanning many centuries and involving the work of multiple authors and editors. In this study, a fully automated method of...
| Publicado en: | Digital Scholarship in the Humanities Vol. 33; no. 3; pp. 670 - 685 |
|---|---|
| Autor principal: | |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
Sep2018
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=131417032&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 131417032 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 2055768X JEO9 jtl: Digital Scholarship in the Humanities issn: 2055768X maglogo: N pubinfo: dt: Sep2018 vid: 33 iid: 3 pid: 622 pub: Oxford University Press / USA artinfo: ui: 131417032 10.1093/llc/fqx024 ppf: 670 ppct: 15 formats: fmt: – @attributes: type: T – @attributes: type: P size: 969KB tig: atl: Unsupervised identification of text reuse in early Chinese literature. aug: au: Sturgeon, Donald affil: Fairbank Center for Chinese Studies, Harvard University, USA su: Chinese literature Online databases Data mining Corpora Stemming (Linguistics) sug: subj: Chinese literature Online databases Data mining Corpora Stemming (Linguistics) ab: Text reuse in early Chinese transmitted texts is extensive and widespread, often reflecting complex textual histories involving repeated transcription, compilation, and editing spanning many centuries and involving the work of multiple authors and editors. In this study, a fully automated method of identifying and representing complex text reuse patterns is presented, and the results evaluated by comparison to a manually compiled reference work. The resultant data are integrated into a widely used and publicly available online database system with browse, search, and visualization functionality. These same results are then aggregated to create a model of text reuse relationships at a corpus level, revealing patterns of systematic reuse among groups of texts. Lastly, the large number of reuse instances identified make possible the analysis of frequently observed string substitutions, which are observed to be strongly indicative of partial synonymy between strings. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Digital Scholarship in the Humanities holder: Oxford University Press / USA dt: @attributes: year: 2018 holdings: @attributes: islocal: N |
|---|