Apache Spark: A Unified Engine for Big Data Processing.

The article discusses the open source computing framework, Apache Spark, which unifies streaming, batch, and interactive big data workloads to unlock new applications. Topics include Spark's use of an RDD programming model, the use of Spark in diverse applications such as batch processing and image...

Descripción completa

Detalles Bibliográficos
Publicado en:Communications of the ACM Vol. 59; no. 11; pp. 56 - 66
Autores principales: ZAHARIA, MATEI, XIN, REYNOLD S., WENDELL, PATRICK, DAS, TATHAGATA, ARMBRUST, MICHAEL, DAVE, ANKUR, XIANGRUI MENG, ROSEN, JOSH, VENKATARAMAN, SHIVARAM, FRANKLIN, MICHAEL J., GHODSI, ALI, GONZALEZ, JOSEPH, SHENKER, SCOTT, STOICA, ION
Formato: Artículo
Publicado: Association for Computing Machinery Nov2016
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=119379442&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 119379442
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        00010782
        ACM
      jtl: Communications of the ACM
      issn: 00010782
      maglogo: N
    pubinfo:
      dt: Nov2016
      vid: 59
      iid: 11
      pid: 68
      pub: Association for Computing Machinery
    artinfo:
      ui:
        119379442
        10.1145/2934664
      ppf: 56
      ppct: 10
      formats:
      tig:
        atl: Apache Spark: A Unified Engine for Big Data Processing.
      aug:
        au:
          ZAHARIA, MATEI
          XIN, REYNOLD S.
          WENDELL, PATRICK
          DAS, TATHAGATA
          ARMBRUST, MICHAEL
          DAVE, ANKUR
          XIANGRUI MENG
          ROSEN, JOSH
          VENKATARAMAN, SHIVARAM
          FRANKLIN, MICHAEL J.
          GHODSI, ALI
          GONZALEZ, JOSEPH
          SHENKER, SCOTT
          STOICA, ION
        affil:
          Assistant professor of computer science at Stanford University, Stanford, CA, and CTO of Databricks, San Francisco, CA
          Chief architect on the Spark team at Databricks, San Francisco, CA
          Vice president of engineering at Databricks, San Francisco, CA
          Software engineer at Databricks, San Francisco, CA
          Graduate student in the Real-Time, Intelligent and Secure Systems Lab at the University of California, Berkeley
          Ph.D. student in the AMPLab at the University of California, Berkeley
          The Liew Family Chair of Computer Science at the University of Chicago
          Director of the AMPLab at the University of California, Berkeley
          The CEO of Databricks and adjunct faculty at the University of California, Berkeley
          Assistant professor in EECS at the University of California, Berkeley
          Professor in EECS at the University of California, Berkeley
          Professor in EECS and co-director of the AMPLab at the University of California, Berkeley
      su:
        Big data
        Open source products
        Software frameworks
        Distributed computing
        Image processing software
      sug:
        subj:
          Big data
          Open source products
          Software frameworks
          Distributed computing
          Image processing software
      ab: The article discusses the open source computing framework, Apache Spark, which unifies streaming, batch, and interactive big data workloads to unlock new applications. Topics include Spark's use of an RDD programming model, the use of Spark in diverse applications such as batch processing and image processing, and the additional cost of Spark over other specialized systems due to fault tolerance.
      pubtype: Periodical
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      dt:
        @attributes:
          year: 2016
    holdings:
      @attributes:
        islocal: N