Efficiently Processing and Storing Library Linked Data using Apache Spark and Parquet.
Resource Description Framework (RDF) is a commonly used data model in the Semantic Web environment. Libraries and various other communities have been using the RDF data model to store valuable data after it is extracted from traditional storage systems. However, because of the large volume of the da...
| Published in: | Information Technology & Libraries Vol. 37; no. 3; pp. 29 - 50 |
|---|---|
| Main Authors: | , , |
| Format: | pictorial research tables/charts Journal Article |
| Published: |
American Library Association
Sep2018
|
| Online Access: | View this record in EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=132055042&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 132055042 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 07309295 ITL jtl: Information Technology & Libraries issn: 07309295 maglogo: N pubinfo: dt: Sep2018 vid: 37 iid: 3 pid: 55 pub: American Library Association place: Chicago, Illinois artinfo: ui: 132055042 132055042 132055042 10.6017/ital.v37i3.10177 132055042 ppf: 29 ppct: 21 formats: fmt: @attributes: type: P tig: atl: Efficiently Processing and Storing Library Linked Data using Apache Spark and Parquet. aug: au: Sharma, Kumar Marjit, Ujjal Biswas, Utpal affil: Research Scholar, Department of Computer Science and Engineering, the University of Kalyani, India sug: subj: Libraries Information Technology Information Storage Methods Data Management Evaluation Data Analytics Evaluation Software Computers and Computerization Programming Languages ab: Resource Description Framework (RDF) is a commonly used data model in the Semantic Web environment. Libraries and various other communities have been using the RDF data model to store valuable data after it is extracted from traditional storage systems. However, because of the large volume of the data, processing and storing it is becoming a nightmare for traditional data- management tools. This challenge demands a scalable and distributed system that can manage data in parallel. In this article, a distributed solution is proposed for efficiently processing and storing the large volume of library linked data stored in traditional storage systems. Apache Spark is used for parallel processing of large data sets and a column-oriented schema is proposed for storing RDF data. The storage system is built on top of Hadoop Distributed File Systems (HDFS) and uses the Apache Parquet format to store data in a compressed form. The experimental evaluation showed that storage requirements were reduced significantly as compared to Jena TDB, Sesame, RDF/XML, and N-Triples file formats. SPARQL queries are processed using Spark SQL to query the compressed data. The experimental evaluation showed a good query response time, which significantly reduces as the number of worker nodes increases. pubtype: Academic Journal doctype: pictorial research tables/charts Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|