NoDB: Efficient Query Execution on Raw Data Files.

As data collections become larger and larger, users are faced with increasing bottlenecks in their data analysis. More data means more time to prepare and to load the data into the database before executing the desired queries. Many applications already avoid using database systems, for example, sci...

Descripción completa

Detalles Bibliográficos
Publicado en:Communications of the ACM Vol. 58; no. 12; pp. 112 - 122
Autores principales: Alagiannis, loannis, Borovica-Gajic, Renata, Branco, Miguel, Idreos, Stratos, Ailamaki, Anastasia
Formato: Artículo
Publicado: Association for Computing Machinery Dec2015
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=111186000&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 111186000
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        00010782
        ACM
      jtl: Communications of the ACM
      issn: 00010782
      maglogo: N
    pubinfo:
      dt: Dec2015
      vid: 58
      iid: 12
      pid: 68
      pub: Association for Computing Machinery
    artinfo:
      ui:
        111186000
        10.1145/2830508
      ppf: 112
      ppct: 10
      formats:
      tig:
        atl: NoDB: Efficient Query Execution on Raw Data Files.
      aug:
        au:
          Alagiannis, loannis
          Borovica-Gajic, Renata
          Branco, Miguel
          Idreos, Stratos
          Ailamaki, Anastasia
        affil:
          École Polytechnique Fédérale de Lausanne, Lausanne, Switzerland.
          Harvard University, Cambridge, MA.
      su:
        Database management
        Technological innovations
        Database management software
        Query languages (Computer science)
        Data analysis software
        Social network analysis
      sug:
        subj:
          Database management
          Technological innovations
          Database management software
          Query languages (Computer science)
          Data analysis software
          Social network analysis
      ab: As data collections become larger and larger, users are faced with increasing bottlenecks in their data analysis. More data means more time to prepare and to load the data into the database before executing the desired queries. Many applications already avoid using database systems, for example, scientific data analysis and social networks, due to the complexity and the increased data-to-query time, that is, the time between getting the data and retrieving its first useful results. For many applications data collections keep growing fast, even on a daily basis, and this data deluge will only increase in the future, where it is expected to have much more data than what we can move or store, let alone analyze. We here present the design and roadmap of a new paradigm in database systems, called NoDB, which do not require data loading while still maintaining the whole feature set of a modern database system. In particular, we show how to make raw data files a first-class citizen, fully integrated with the query engine. Through our design and lessons learned by implementing the NoDB philosophy over a modern Database Management Systems (DBMS), we discuss the fundamental limitations as well as the strong opportunities that such a research path brings. We identify performance bottlenecks specific for in situ processing, namely the repeated parsing and tokenizing overhead and the expensive data type conversion. To address these problems, we introduce an adaptive indexing mechanism that maintains positional information to provide efficient access to raw data files, together with a flexible caching structure. We conclude that NoDB systems are feasible to design and implement over modern DBMS, bringing an unprecedented positive effect in usability and performance.
      pubtype: Periodical
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      dt:
        @attributes:
          year: 2015
    holdings:
      @attributes:
        islocal: N