Exploring Finnic written oral folk poetry through string similarity.
Suomen Kansan Vanhat Runot (Old Poems of the Finnish People) is a collection of nearly 90,000 oral folk poems written down between 1564 and the early 20th century. It is characterized by frequent reoccurrence of similar pieces of text on various levels (from entire poems, through passages to single...
| Publicado en: | Digital Scholarship in the Humanities Vol. 38; no. 1; pp. 180 - 195 |
|---|---|
| Autores principales: | , , |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
Apr2023
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=162941097&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 162941097 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 2055768X JEO9 jtl: Digital Scholarship in the Humanities issn: 2055768X maglogo: N pubinfo: dt: Apr2023 vid: 38 iid: 1 pid: 622 pub: Oxford University Press / USA artinfo: ui: 162941097 10.1093/llc/fqac034 ppf: 180 ppct: 15 formats: fmt: – @attributes: type: T – @attributes: type: P size: 865KB tig: atl: Exploring Finnic written oral folk poetry through string similarity. aug: au: Janicki, Maciej Kallio, Kati Sarv, Mari affil: Department of Digital Humanities, University of Helsinki , Finland Finnish Literature Society , Finland Estonian Literary Museum , Estonia su: Semantics Web browsing Poetry (Literary form) Poetry writing Twentieth century sug: subj: Semantics Web browsing Poetry (Literary form) Poetry writing Twentieth century ab: Suomen Kansan Vanhat Runot (Old Poems of the Finnish People) is a collection of nearly 90,000 oral folk poems written down between 1564 and the early 20th century. It is characterized by frequent reoccurrence of similar pieces of text on various levels (from entire poems, through passages to single verses and collocations). However, finding these similarities is challenging due to a high degree of orthographical, morphological, and compositional variation. In this article, we propose a method for automatically identifying equivalent verses , i.e. verses conveying the same meaning with the same words, using a clustering based on cosine similarity of character bigram vectors. The method achieves around 81% F-score and has been successfully used for identifying similarities across the entire SKVR corpus on the level of verse, passage, and poem. The results can be browsed through a Web interface. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Digital Scholarship in the Humanities holder: Oxford University Press / USA dt: @attributes: year: 2023 holdings: @attributes: islocal: N |
|---|