Apache Spark: A Unified Engine for Big Data Processing.
The article discusses the open source computing framework, Apache Spark, which unifies streaming, batch, and interactive big data workloads to unlock new applications. Topics include Spark's use of an RDD programming model, the use of Spark in diverse applications such as batch processing and image...
| Publicado en: | Communications of the ACM Vol. 59; no. 11; pp. 56 - 66 |
|---|---|
| Autores principales: | , , , , , , , , , , , , , |
| Formato: | Artículo |
| Publicado: |
Association for Computing Machinery
Nov2016
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=119379442&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 119379442 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 00010782 ACM jtl: Communications of the ACM issn: 00010782 maglogo: N pubinfo: dt: Nov2016 vid: 59 iid: 11 pid: 68 pub: Association for Computing Machinery artinfo: ui: 119379442 10.1145/2934664 ppf: 56 ppct: 10 formats: tig: atl: Apache Spark: A Unified Engine for Big Data Processing. aug: au: ZAHARIA, MATEI XIN, REYNOLD S. WENDELL, PATRICK DAS, TATHAGATA ARMBRUST, MICHAEL DAVE, ANKUR XIANGRUI MENG ROSEN, JOSH VENKATARAMAN, SHIVARAM FRANKLIN, MICHAEL J. GHODSI, ALI GONZALEZ, JOSEPH SHENKER, SCOTT STOICA, ION affil: Assistant professor of computer science at Stanford University, Stanford, CA, and CTO of Databricks, San Francisco, CA Chief architect on the Spark team at Databricks, San Francisco, CA Vice president of engineering at Databricks, San Francisco, CA Software engineer at Databricks, San Francisco, CA Graduate student in the Real-Time, Intelligent and Secure Systems Lab at the University of California, Berkeley Ph.D. student in the AMPLab at the University of California, Berkeley The Liew Family Chair of Computer Science at the University of Chicago Director of the AMPLab at the University of California, Berkeley The CEO of Databricks and adjunct faculty at the University of California, Berkeley Assistant professor in EECS at the University of California, Berkeley Professor in EECS at the University of California, Berkeley Professor in EECS and co-director of the AMPLab at the University of California, Berkeley su: Big data Open source products Software frameworks Distributed computing Image processing software sug: subj: Big data Open source products Software frameworks Distributed computing Image processing software ab: The article discusses the open source computing framework, Apache Spark, which unifies streaming, batch, and interactive big data workloads to unlock new applications. Topics include Spark's use of an RDD programming model, the use of Spark in diverse applications such as batch processing and image processing, and the additional cost of Spark over other specialized systems due to fault tolerance. pubtype: Periodical doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y dt: @attributes: year: 2016 holdings: @attributes: islocal: N |
|---|