Convolution Engine: Balancing Efficiency and Flexibility in Specialized Computing.
General-purpose processors, while tremendously versatile, pay a huge cost for their flexibility by wasting over 99% of the energy in programmability overheads. We observe that reducing this waste requires tuning data storage and compute structures and their connectivity to the data-flow and data-loc...
| Publicado en: | Communications of the ACM Vol. 58; no. 4; pp. 85 - 94 |
|---|---|
| Autores principales: | , , , , , |
| Formato: | Artículo |
| Publicado: |
Association for Computing Machinery
Apr2015
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=101826068&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 101826068 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 00010782 ACM jtl: Communications of the ACM issn: 00010782 maglogo: N pubinfo: dt: Apr2015 vid: 58 iid: 4 pid: 68 pub: Association for Computing Machinery artinfo: ui: 101826068 10.1145/2735841 ppf: 85 ppct: 9 formats: tig: atl: Convolution Engine: Balancing Efficiency and Flexibility in Specialized Computing. aug: au: Qadeer, Wajahat Hameed, Rehan Shacham, Ofer Venkatesan, Preethi Kozyrakis, Christos Horowitz, Mark affil: Google, Mountain View, CA Intel Corporation, Santa Clara, CA Stanford University, Stanford, CA su: Microprocessors Information retrieval Data warehousing Signal convolution Flow control (Data transmission systems) Image processing sug: subj: Microprocessors Information retrieval Data warehousing Signal convolution Flow control (Data transmission systems) Image processing ab: General-purpose processors, while tremendously versatile, pay a huge cost for their flexibility by wasting over 99% of the energy in programmability overheads. We observe that reducing this waste requires tuning data storage and compute structures and their connectivity to the data-flow and data-locality patterns in the algorithms. Hence, by backing off from full programmability and instead targeting key data-flow patterns used in a domain, we can create efficient engines that can be programmed and reused across a wide range of applications within that domain. We present the Convolution Engine (CE)--a programmable processor specialized for the convolution-like data-flow prevalent in computational photography, computer vision, and video processing. The CE achieves energy efficiency by capturing data-reuse patterns, eliminating data transfer overheads, and enabling a large number of operations per memory access. We demonstrate that the CE is within a factor of 2–3× of the energy and area efficiency of custom units optimized for a single kernel. The CE improves energy and area efficiency by 8–15× over data-parallel Single Instruction Multiple Data (SIMD) engines for most image processing applications. pubtype: Periodical doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y dt: @attributes: year: 2015 holdings: @attributes: islocal: N |
|---|