Extracting structured data from publications in the Art Conservation Domain.
The most common method of publishing new discoveries about art conservation techniques and research has been through traditional full-text publications. Such corpora typically only support searching via metadata (e.g. title, authors, or keywords) and full-text. In particular, it is difficult to disc...
| Publicado en: | Digital Scholarship in the Humanities pp. 225 - 246 |
|---|---|
| Autores principales: | , , |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
06/01/2015
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=108489947&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 108489947 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 2055768X JEO9 jtl: Digital Scholarship in the Humanities issn: 2055768X maglogo: N pubinfo: dt: 06/01/2015 pid: 622 pub: Oxford University Press / USA artinfo: ui: 108489947 10.1093/llc/fqu002 ppf: 225 ppct: 21 formats: fmt: @attributes: type: P size: 10.3MB tig: atl: Extracting structured data from publications in the Art Conservation Domain. aug: au: Odat, Suleiman Groza, Tudor Hunter, Jane affil: School of ITEE, University of Queensland, Australia su: Art conservation & restoration Extraction (Linguistics) Information storage & retrieval systems -- Arts Computers in the arts Machine learning Metadata sug: subj: Art conservation & restoration Extraction (Linguistics) Information storage & retrieval systems -- Arts Computers in the arts Machine learning Metadata ab: The most common method of publishing new discoveries about art conservation techniques and research has been through traditional full-text publications. Such corpora typically only support searching via metadata (e.g. title, authors, or keywords) and full-text. In particular, it is difficult to discover valuable information about the chemical processes, experimental results, or preservation treatments associated with the conservation of paintings from a specific genre. This article addresses this problem by focusing on the extraction of structured data (that complies with a pre-defined ontology) from a distributed corpus of publications about painting conservation. Our specific extraction method involves a unique combination of named entity recognition (using gazetteer-based and machine learning-based methods) followed by relationship extraction (using rule-based and machine learning-based methods). The resulting structured data are stored in a resource description framework triple store, and a Web-based graphical user interface enables the SPARQL querying, retrieval, and display of the search results. The results from applying our techniques to a corpus of publications on art conservation indicate that our approach achieves higher quality precision and recall in extracting named entities and relations from publications, relative to alternative existing approaches. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Digital Scholarship in the Humanities holder: Oxford University Press / USA dt: @attributes: year: 2015 holdings: @attributes: islocal: N |
|---|