HammingMesh: A Network Topology for Large-Scale Deep Learning.

This article presents HammingMesh, a flexible topology that overcomes current high-performance computing’s inability to support deep-learning workloads by allowing for the adjustment of the ratio of local and global bandwidth. The article discusses communication in distributed deep learning with a l...

Descripción completa

Detalles Bibliográficos
Publicado en:Communications of the ACM Vol. 67; no. 12; pp. 97 - 106
Autores principales: Hoefler, Torsten, Bonoto, Tommaso, De Sensi, Daniele, Di Girolamo, Salvatore, Li, Shigang, Heddes, Marco, Goel, Deepak, Castro, Miguel, Scott, Steve
Formato: Artículo
Publicado: Association for Computing Machinery Dec2024
Materias:
Acceso en línea:Ver este registro en EBSCOhost
Descripción
Sumario:This article presents HammingMesh, a flexible topology that overcomes current high-performance computing’s inability to support deep-learning workloads by allowing for the adjustment of the ratio of local and global bandwidth. The article discusses communication in distributed deep learning with a look at data parallelism, pipeline parallelism, and operator parallelism. Topics include both bisection and global bandwidth, logical job topologies and failures, and microbenchmarks in HammingMesh.