HammingMesh: A Network Topology for Large-Scale Deep Learning.

This article presents HammingMesh, a flexible topology that overcomes current high-performance computing’s inability to support deep-learning workloads by allowing for the adjustment of the ratio of local and global bandwidth. The article discusses communication in distributed deep learning with a l...

Descripción completa

Detalles Bibliográficos
Publicado en:Communications of the ACM Vol. 67; no. 12; pp. 97 - 106
Autores principales: Hoefler, Torsten, Bonoto, Tommaso, De Sensi, Daniele, Di Girolamo, Salvatore, Li, Shigang, Heddes, Marco, Goel, Deepak, Castro, Miguel, Scott, Steve
Formato: Artículo
Publicado: Association for Computing Machinery Dec2024
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=181072021&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 181072021
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        00010782
        ACM
      jtl: Communications of the ACM
      issn: 00010782
      maglogo: N
    pubinfo:
      dt: Dec2024
      vid: 67
      iid: 12
      pid: 68
      pub: Association for Computing Machinery
    artinfo:
      ui:
        181072021
        10.1145/3623490
      ppf: 97
      ppct: 9
      formats:
      tig:
        atl: HammingMesh: A Network Topology for Large-Scale Deep Learning.
      aug:
        au:
          Hoefler, Torsten
          Bonoto, Tommaso
          De Sensi, Daniele
          Di Girolamo, Salvatore
          Li, Shigang
          Heddes, Marco
          Goel, Deepak
          Castro, Miguel
          Scott, Steve
        affil:
          Microsoft Corp., Zurich, Switzerland
          ETH Zurich, Zurich, Switzerland
          Microsoft Corp, Redmond, WA, USA
          Microsoft Corp, Sunnyvale, CA, USA
          Microsoft Corp, Cambridge, United Kingdom
      su:
        Deep learning
        High performance computing
        Computer networks
        Bandwidth allocation
        Artificial neural networks
      sug:
        subj:
          Deep learning
          High performance computing
          Computer networks
          Bandwidth allocation
          Artificial neural networks
      ab: This article presents HammingMesh, a flexible topology that overcomes current high-performance computing’s inability to support deep-learning workloads by allowing for the adjustment of the ratio of local and global bandwidth. The article discusses communication in distributed deep learning with a look at data parallelism, pipeline parallelism, and operator parallelism. Topics include both bisection and global bandwidth, logical job topologies and failures, and microbenchmarks in HammingMesh.
      pubtype: Periodical
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      dt:
        @attributes:
          year: 2024
    holdings:
      @attributes:
        islocal: N