Extracting structured data from publications in the Art Conservation Domain.

The most common method of publishing new discoveries about art conservation techniques and research has been through traditional full-text publications. Such corpora typically only support searching via metadata (e.g. title, authors, or keywords) and full-text. In particular, it is difficult to disc...

Descripción completa

Detalles Bibliográficos
Publicado en:Digital Scholarship in the Humanities pp. 225 - 246
Autores principales: Odat, Suleiman, Groza, Tudor, Hunter, Jane
Formato: Artículo
Publicado: Oxford University Press / USA 06/01/2015
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=108489947&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 108489947
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        2055768X
        JEO9
      jtl: Digital Scholarship in the Humanities
      issn: 2055768X
      maglogo: N
    pubinfo:
      dt: 06/01/2015
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        108489947
        10.1093/llc/fqu002
      ppf: 225
      ppct: 21
      formats:
        fmt:
          @attributes:
            type: P
            size: 10.3MB
      tig:
        atl: Extracting structured data from publications in the Art Conservation Domain.
      aug:
        au:
          Odat, Suleiman
          Groza, Tudor
          Hunter, Jane
        affil: School of ITEE, University of Queensland, Australia
      su:
        Art conservation & restoration
        Extraction (Linguistics)
        Information storage & retrieval systems -- Arts
        Computers in the arts
        Machine learning
        Metadata
      sug:
        subj:
          Art conservation & restoration
          Extraction (Linguistics)
          Information storage & retrieval systems -- Arts
          Computers in the arts
          Machine learning
          Metadata
      ab: The most common method of publishing new discoveries about art conservation techniques and research has been through traditional full-text publications. Such corpora typically only support searching via metadata (e.g. title, authors, or keywords) and full-text. In particular, it is difficult to discover valuable information about the chemical processes, experimental results, or preservation treatments associated with the conservation of paintings from a specific genre. This article addresses this problem by focusing on the extraction of structured data (that complies with a pre-defined ontology) from a distributed corpus of publications about painting conservation. Our specific extraction method involves a unique combination of named entity recognition (using gazetteer-based and machine learning-based methods) followed by relationship extraction (using rule-based and machine learning-based methods). The resulting structured data are stored in a resource description framework triple store, and a Web-based graphical user interface enables the SPARQL querying, retrieval, and display of the search results. The results from applying our techniques to a corpus of publications on art conservation indicate that our approach achieves higher quality precision and recall in extracting named entities and relations from publications, relative to alternative existing approaches.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Digital Scholarship in the Humanities
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2015
    holdings:
      @attributes:
        islocal: N