Enabling complex analysis of large-scale digital collections: humanities research, high-performance computing, and transforming access to British Library digital collections.

Although there has been a drive in the cultural heritage sector to provide largescale, open data sets for researchers, we have not seen a commensurate rise in humanities researchers undertaking complex analysis of these data sets for their own research purposes. This article reports on a pilot proje...

Full description

Bibliographic Details
Published in:Digital Scholarship in the Humanities Vol. 33; no. 2; pp. 456 - 467
Main Authors: Terras, Melissa, Baker, James, Hetherington, James, Beavan, David, Austwick, Martin Zaltz, Welsh, Anne, O'Neill, Helen, Finley, Will, Duke-Williams, Oliver, Farquhar, Adam
Format: Article
Published: Oxford University Press / USA Jun2018
Subjects:
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=129666132&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 129666132
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        2055768X
        JEO9
      jtl: Digital Scholarship in the Humanities
      issn: 2055768X
      maglogo: N
    pubinfo:
      dt: Jun2018
      vid: 33
      iid: 2
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        129666132
        10.1093/llc/fqx020
      ppf: 456
      ppct: 11
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 602KB
      tig:
        atl: Enabling complex analysis of large-scale digital collections: humanities research, high-performance computing, and transforming access to British Library digital collections.
      aug:
        au:
          Terras, Melissa
          Baker, James
          Hetherington, James
          Beavan, David
          Austwick, Martin Zaltz
          Welsh, Anne
          O'Neill, Helen
          Finley, Will
          Duke-Williams, Oliver
          Farquhar, Adam
        affil:
          Department of Information Studies, University College London, UK and UCL Centre for Digital Humanities, University College London, UK
          School of History, Art History and Philosophy, University of Sussex, UK
          Research Software Development Group, Research IT Services, University College London, UK
          UCL Centre for Digital Humanities, University College London, UK
          Centre for Advanced Spatial Analysis, University College London, UK
          Department of Information Studies, University College London, UK
          Department of Information Studies, University College London, UK, The London Library, UK
          Department of History, University of Sheffield, UK
          Digital Scholarship, British Library, UK
      su:
        Cultural property
        Humanities
        University College, London
        British Library
        Academic-industrial collaboration
      sug:
        subj:
          Cultural property
          Humanities
          University College, London
          British Library
          Academic-industrial collaboration
      ab: Although there has been a drive in the cultural heritage sector to provide largescale, open data sets for researchers, we have not seen a commensurate rise in humanities researchers undertaking complex analysis of these data sets for their own research purposes. This article reports on a pilot project at University College London, working in collaboration with the British Library, to scope out how best high-performance computing facilities can be used to facilitate the needs of researchers in the humanities. Using institutional data-processing frameworks routinely used to support scientific research, we assisted four humanities researchers in analysing 60,000 digitized books, and we present two resulting case studies here. This research allowed us to identify infrastructural and procedural barriers and make recommendations on resource allocation to best support non-computational researchers in undertaking 'big data' research. We recommend that research software engineer capacity can be most efficiently deployed in maintaining and supporting data sets, while librarians can provide an essential service in running initial, routine queries for humanities scholars. At present there are too many technical hurdles for most individuals in the humanities to consider analysing at scale these increasingly available open data sets, and by building on existing frameworks of support from research computing and library services, we can best support humanities scholars in developing methods and approaches to take advantage of these research opportunities.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Digital Scholarship in the Humanities
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2018
    holdings:
      @attributes:
        islocal: N