HammingMesh: A Network Topology for Large-Scale Deep Learning.
This article presents HammingMesh, a flexible topology that overcomes current high-performance computing’s inability to support deep-learning workloads by allowing for the adjustment of the ratio of local and global bandwidth. The article discusses communication in distributed deep learning with a l...
| Publicado en: | Communications of the ACM Vol. 67; no. 12; pp. 97 - 106 |
|---|---|
| Autores principales: | , , , , , , , , |
| Formato: | Artículo |
| Publicado: |
Association for Computing Machinery
Dec2024
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| Sumario: | This article presents HammingMesh, a flexible topology that overcomes current high-performance computing’s inability to support deep-learning workloads by allowing for the adjustment of the ratio of local and global bandwidth. The article discusses communication in distributed deep learning with a look at data parallelism, pipeline parallelism, and operator parallelism. Topics include both bisection and global bandwidth, logical job topologies and failures, and microbenchmarks in HammingMesh. |
|---|