A Massively Parallel Adaptive Fast Multipole Method on Heterogeneous Architectures.
We describe a parallel fast multipole method (FMM) for highly nonuniform distributions of particles. We employ both distributed memory parallelism (via MPI) and shared memory parallelism (via OpenMP and GPU acceleration) to rapidly evaluate two-body nonoscillatory potentials in three dimensions on h...
| Publicado en: | Communications of the ACM Vol. 55; no. 5; pp. 101 - 110 |
|---|---|
| Formato: | Artículo |
| Publicado: |
Association for Computing Machinery
May2012
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=74715569&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 74715569 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 00010782 ACM jtl: Communications of the ACM issn: 00010782 maglogo: N pubinfo: dt: May2012 vid: 55 iid: 5 pid: 68 pub: Association for Computing Machinery artinfo: ui: 74715569 10.1145/2160718.2160740 ppf: 101 ppct: 9 formats: tig: atl: A Massively Parallel Adaptive Fast Multipole Method on Heterogeneous Architectures. aug: su: Particle size distribution Parallel processing Heterogeneity Computer architecture Scalability Graphics processing units Multicore processors sug: subj: Particle size distribution Parallel processing Heterogeneity Computer architecture Scalability Graphics processing units Multicore processors ab: We describe a parallel fast multipole method (FMM) for highly nonuniform distributions of particles. We employ both distributed memory parallelism (via MPI) and shared memory parallelism (via OpenMP and GPU acceleration) to rapidly evaluate two-body nonoscillatory potentials in three dimensions on heterogeneous high performance computing architectures. We have performed scalability tests with up to 30 billion particles on 196,608 cores on the AMD/CRAY-based Jaguar system at ORNL. On a GPU-enabled system (NSF’s Keeneland at Georgia Tech/ORNL), we observed 30× speedup over a single core CPU and 7× speedup over a multicore CPU implementation. By combining GPUs with MPI, we achieve less than 10 ns/particle and six digits of accuracy for a run with 48 million nonuniformly distributed particles on 192 GPUs. pubtype: Periodical doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y dt: @attributes: year: 2012 holdings: @attributes: islocal: N |
|---|