Learning Your Limit: Managing Massively Multithreaded Caches Through Scheduling.
The gap between processor and memory performance has become a focal point for microprocessor research and development over the past three decades. Modern architectures use two orthogonal approaches to help alleviate this issue: (1) Almost every microprocessor includes some form of on-chip storage, u...
| Publicado en: | Communications of the ACM Vol. 57; no. 12; pp. 91 - 99 |
|---|---|
| Autores principales: | , , |
| Formato: | Artículo |
| Publicado: |
Association for Computing Machinery
Dec2014
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=99744615&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 99744615 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 00010782 ACM jtl: Communications of the ACM issn: 00010782 maglogo: N pubinfo: dt: Dec2014 vid: 57 iid: 12 pid: 68 pub: Association for Computing Machinery artinfo: ui: 99744615 10.1145/2682583 ppf: 91 ppct: 8 formats: tig: atl: Learning Your Limit: Managing Massively Multithreaded Caches Through Scheduling. aug: au: Rogers, Timothy G. O’Connor, Mike Aamodt, Tor M. affil: Department of Electrical and Computer Engineering, University of British Columbia, Vancouver, Canada NVIDIA Research, Austin, TX su: Cache memory Computer storage devices Microprocessors Computer scheduling Computer architecture Computer science sug: subj: Cache memory Computer storage devices Microprocessors Computer scheduling Computer architecture Computer science ab: The gap between processor and memory performance has become a focal point for microprocessor research and development over the past three decades. Modern architectures use two orthogonal approaches to help alleviate this issue: (1) Almost every microprocessor includes some form of on-chip storage, usually in the form of caches, to decrease memory latency and make more effective use of limited memory bandwidth. (2) Massively multithreaded architectures, such as graphics processing units (GPUs), attempt to hide the high latency to memory by rapidly switching between many threads directly in hardware. This paper explores the intersection of these two techniques. We study the effect of accelerating highly parallel workloads with significant locality on a massively multithreaded GPU. We observe that the memory access stream seen by on-chip caches is the direct result of decisions made by the hardware thread scheduler. Our work proposes a hardware scheduling technique that reacts to feedback from the memory system to create a more cache-friendly access stream. We evaluate our technique using simulations and show a significant performance improvement over previously proposed scheduling mechanisms. We demonstrate the effectiveness of scheduling as a cache management technique by comparing cache hit rate using our scheduler and an LRU replacement policy against other scheduling techniques using an optimal cache replacement policy. pubtype: Periodical doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y dt: @attributes: year: 2014 holdings: @attributes: islocal: N |
|---|