HammingMesh: A Network Topology for Large-Scale Deep Learning.
This article presents HammingMesh, a flexible topology that overcomes current high-performance computing’s inability to support deep-learning workloads by allowing for the adjustment of the ratio of local and global bandwidth. The article discusses communication in distributed deep learning with a l...
| Publicado en: | Communications of the ACM Vol. 67; no. 12; pp. 97 - 106 |
|---|---|
| Autores principales: | , , , , , , , , |
| Formato: | Artículo |
| Publicado: |
Association for Computing Machinery
Dec2024
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=181072021&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 181072021 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 00010782 ACM jtl: Communications of the ACM issn: 00010782 maglogo: N pubinfo: dt: Dec2024 vid: 67 iid: 12 pid: 68 pub: Association for Computing Machinery artinfo: ui: 181072021 10.1145/3623490 ppf: 97 ppct: 9 formats: tig: atl: HammingMesh: A Network Topology for Large-Scale Deep Learning. aug: au: Hoefler, Torsten Bonoto, Tommaso De Sensi, Daniele Di Girolamo, Salvatore Li, Shigang Heddes, Marco Goel, Deepak Castro, Miguel Scott, Steve affil: Microsoft Corp., Zurich, Switzerland ETH Zurich, Zurich, Switzerland Microsoft Corp, Redmond, WA, USA Microsoft Corp, Sunnyvale, CA, USA Microsoft Corp, Cambridge, United Kingdom su: Deep learning High performance computing Computer networks Bandwidth allocation Artificial neural networks sug: subj: Deep learning High performance computing Computer networks Bandwidth allocation Artificial neural networks ab: This article presents HammingMesh, a flexible topology that overcomes current high-performance computing’s inability to support deep-learning workloads by allowing for the adjustment of the ratio of local and global bandwidth. The article discusses communication in distributed deep learning with a look at data parallelism, pipeline parallelism, and operator parallelism. Topics include both bisection and global bandwidth, logical job topologies and failures, and microbenchmarks in HammingMesh. pubtype: Periodical doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y dt: @attributes: year: 2024 holdings: @attributes: islocal: N |
|---|