Exploring entity recognition and disambiguation for cultural heritage collections.
Unstructured metadata fields such as 'description' offer tremendous value for users to understand cultural heritage objects. However, this type of narrative information is of little direct use within a machine-readable context due to its unstructured nature. This article explores the possibilities a...
| Publicado en: | Digital Scholarship in the Humanities pp. 262 - 280 |
|---|---|
| Autores principales: | , , , , |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
06/01/2015
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=108489949&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 108489949 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 2055768X JEO9 jtl: Digital Scholarship in the Humanities issn: 2055768X maglogo: N pubinfo: dt: 06/01/2015 pid: 622 pub: Oxford University Press / USA artinfo: ui: 108489949 10.1093/llc/fqt067 ppf: 262 ppct: 18 formats: fmt: @attributes: type: P size: 9.2MB tig: atl: Exploring entity recognition and disambiguation for cultural heritage collections. aug: au: van Hooland, Seth De Wilde, Max Verborgh, Ruben Steiner, Thomas Van de Walle, Rik affil: Université libre de Bruxelles (ULB), Brussels, Belgium iMinds--Multimedia LabGhent University, Ghent, Belgium Universitat Politécnica de Catalunya, Barcelona, Spain su: Cultural property Data mining Digital humanities Electronic records Linked data (Semantic Web) Data transformations (Statistics) sug: subj: Cultural property Data mining Digital humanities Electronic records Linked data (Semantic Web) Data transformations (Statistics) ab: Unstructured metadata fields such as 'description' offer tremendous value for users to understand cultural heritage objects. However, this type of narrative information is of little direct use within a machine-readable context due to its unstructured nature. This article explores the possibilities and limitations of named-entity recognition (NER) and term extraction (TE) to mine such unstructured metadata for meaningful concepts. These concepts can be used to leverage otherwise limited searching and browsing operations, but they can also play an important role to foster Digital Humanities research. To catalyze experimentation with NER and TE, the article proposes an evaluation of the performance of three third-party entity extraction services through a comprehensive case study, based on the descriptive fields of the Smithsonian Cooper-Hewitt National Design Museum in New York. To cover both NER and TE, we first offer a quantitative analysis of named entities retrieved by the services in terms of precision and recall compared with a manually annotated gold-standard corpus, and then complement this approach with a more qualitative assessment of relevant terms extracted. Based on the outcomes of this double analysis, the conclusions present the added value of entity extraction services, but also indicate the dangers of uncritically using NER and/or TE, and by extension Linked Data principles, within the Digital Humanities. All metadata and tools used within the article are freely available, making it possible for researchers and practitioners to repeat the methodology. By doing so, the article offers a significant contribution towards understanding the value of entity recognition and disambiguation for the Digital Humanities. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Digital Scholarship in the Humanities holder: Oxford University Press / USA dt: @attributes: year: 2015 holdings: @attributes: islocal: N |
|---|