Understanding Sources of Inefficiency in General-Purpose Chips.

Scaling the performance of a power limited processor requires decreasing the energy expended per instruction executed, since energy/op * op/second is power. To better understand what improvement in processor efficiency is possible, and what must be done to capture it, we quantify the sources of the...

Full description

Bibliographic Details
Published in:Communications of the ACM Vol. 54; no. 10; pp. 85 - 94
Main Authors: Hameed, Rehan, Qadeer, Wajahat, Wachs, Megan, Azizi, Omid, Solomatnikov, Alex, Lee, Benjamin C., Richardson, Stephen, Kozyrakis, Christos, Horowitz, Mark
Format: Article
Published: Association for Computing Machinery Oct2011
Subjects:
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=66736075&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 66736075
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        00010782
        ACM
      jtl: Communications of the ACM
      issn: 00010782
      maglogo: N
    pubinfo:
      dt: Oct2011
      vid: 54
      iid: 10
      pid: 68
      pub: Association for Computing Machinery
    artinfo:
      ui:
        66736075
        10.1145/2001269.2001291
      ppf: 85
      ppct: 9
      formats:
      tig:
        atl: Understanding Sources of Inefficiency in General-Purpose Chips.
      aug:
        au:
          Hameed, Rehan
          Qadeer, Wajahat
          Wachs, Megan
          Azizi, Omid
          Solomatnikov, Alex
          Lee, Benjamin C.
          Richardson, Stephen
          Kozyrakis, Christos
          Horowitz, Mark
        affil:
          Stanford University, Stanford, CA
          Hicamp Systems, Menlo Park, CA
          Duke University, Durham, NC
      su:
        Integrated circuits
        Central processing units
        Energy consumption
        Algorithm software
        Custom computer software
        Software architecture
      sug:
        subj:
          Integrated circuits
          Central processing units
          Energy consumption
          Algorithm software
          Custom computer software
          Software architecture
      ab: Scaling the performance of a power limited processor requires decreasing the energy expended per instruction executed, since energy/op * op/second is power. To better understand what improvement in processor efficiency is possible, and what must be done to capture it, we quantify the sources of the performance and energy overheads of a 720p HD H.264 encoder running on a general-purpose fourprocessor CMP system. The initial overheads are large: the CMP was 500x less energy efficient than an Application Specific Integrated Circuit (ASIC) doing the same job. We explore methods to eliminate these overheads by transforming the CPU into a specialized system for H.264 encoding. Broadly applicable optimizations like single instruction, multiple data (SIMD) units improve CMP performance by 14x and energy by 10x, which is still 50x worse than an ASIC. The problem is that the basic operation costs in H.264 are so small that even with a SIMD unit doing over 10 ops per cycle, 90% of the energy is still overhead. Achieving ASIC-like performance and efficiency requires algorithm-specific optimizations. For each subalgorithm of H.264, we create a large, specialized functional/storage unit capable of executing hundreds of operations per instruction. This improves energy efficiency by 160x (instead of 10x), and the final customized CMP reaches the same performance and within 3x of an ASIC solution's energy in comparable area.
      pubtype: Periodical
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      dt:
        @attributes:
          year: 2011
    holdings:
      @attributes:
        islocal: N