Efficiently Processing and Storing Library Linked Data using Apache Spark and Parquet.

Resource Description Framework (RDF) is a commonly used data model in the Semantic Web environment. Libraries and various other communities have been using the RDF data model to store valuable data after it is extracted from traditional storage systems. However, because of the large volume of the da...

Full description

Bibliographic Details
Published in:Information Technology & Libraries Vol. 37; no. 3; pp. 29 - 50
Main Authors: Sharma, Kumar, Marjit, Ujjal, Biswas, Utpal
Format: pictorial research tables/charts Journal Article
Published: American Library Association Sep2018
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=132055042&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 132055042
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        07309295
        ITL
      jtl: Information Technology & Libraries
      issn: 07309295
      maglogo: N
    pubinfo:
      dt: Sep2018
      vid: 37
      iid: 3
      pid: 55
      pub: American Library Association
      place: Chicago, Illinois
    artinfo:
      ui:
        132055042
        132055042
        132055042
        10.6017/ital.v37i3.10177
        132055042
      ppf: 29
      ppct: 21
      formats:
        fmt:
          @attributes:
            type: P
      tig:
        atl: Efficiently Processing and Storing Library Linked Data using Apache Spark and Parquet.
      aug:
        au:
          Sharma, Kumar
          Marjit, Ujjal
          Biswas, Utpal
        affil: Research Scholar, Department of Computer Science and Engineering, the University of Kalyani, India
      sug:
        subj:
          Libraries
          Information Technology
          Information Storage Methods
          Data Management Evaluation
          Data Analytics Evaluation
          Software
          Computers and Computerization
          Programming Languages
      ab: Resource Description Framework (RDF) is a commonly used data model in the Semantic Web environment. Libraries and various other communities have been using the RDF data model to store valuable data after it is extracted from traditional storage systems. However, because of the large volume of the data, processing and storing it is becoming a nightmare for traditional data- management tools. This challenge demands a scalable and distributed system that can manage data in parallel. In this article, a distributed solution is proposed for efficiently processing and storing the large volume of library linked data stored in traditional storage systems. Apache Spark is used for parallel processing of large data sets and a column-oriented schema is proposed for storing RDF data. The storage system is built on top of Hadoop Distributed File Systems (HDFS) and uses the Apache Parquet format to store data in a compressed form. The experimental evaluation showed that storage requirements were reduced significantly as compared to Jena TDB, Sesame, RDF/XML, and N-Triples file formats. SPARQL queries are processed using Spark SQL to query the compressed data. The experimental evaluation showed a good query response time, which significantly reduces as the number of worker nodes increases.
      pubtype: Academic Journal
      doctype:
        pictorial
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N