Halide: Decoupling Algorithms from Schedules for High-Performance Image Processing.
Writing high-performance code on modern machines requires not just locally optimizing inner loops, but globally reorganizing computations to exploit parallelism and locality--doing things such as tiling and blocking whole pipelines to fit in cache. This is especially true for image processing pipeli...
| Published in: | Communications of the ACM Vol. 61; no. 1; pp. 106 - 116 |
|---|---|
| Main Authors: | , , , , , , , |
| Format: | Article |
| Published: |
Association for Computing Machinery
Jan2018
|
| Subjects: | |
| Online Access: | View this record in EBSCOhost |