Pivot Tracing: Dynamic Causal Monitoring for Distributed Systems.

Monitoring and troubleshooting distributed systems are notoriously difficult; potential problems are complex, varied, and unpredictable. The monitoring and diagnosis tools commonly used today—logs, counters, and metrics—have two important limitations: what gets recorded is defined a priori, and the...

Descripción completa

Detalles Bibliográficos
Publicado en:Communications of the ACM Vol. 63; no. 3; pp. 94 - 103
Autores principales: Mace, Jonathan, Roelke, Ryan, Fonseca, Rodrigo
Formato: Artículo
Publicado: Association for Computing Machinery Mar2020
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=141926857&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 141926857
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        00010782
        ACM
      jtl: Communications of the ACM
      issn: 00010782
      maglogo: N
    pubinfo:
      dt: Mar2020
      vid: 63
      iid: 3
      pid: 68
      pub: Association for Computing Machinery
    artinfo:
      ui:
        141926857
        10.1145/3378933
      ppf: 94
      ppct: 9
      formats:
      tig:
        atl: Pivot Tracing: Dynamic Causal Monitoring for Distributed Systems.
      aug:
        au:
          Mace, Jonathan
          Roelke, Ryan
          Fonseca, Rodrigo
        affil: Brown University Department of Computer Science, Providence, RI, USA.
      su:
        Computer systems
        Debugging
        Computer science
        Electronic data processing
        Query languages (Computer science)
      sug:
        subj:
          Computer systems
          Debugging
          Computer science
          Electronic data processing
          Query languages (Computer science)
      ab: Monitoring and troubleshooting distributed systems are notoriously difficult; potential problems are complex, varied, and unpredictable. The monitoring and diagnosis tools commonly used today—logs, counters, and metrics—have two important limitations: what gets recorded is defined a priori, and the information is recorded in a component- or machine-centric way, making it extremely hard to correlate events that cross these boundaries. This paper presents Pivot Tracing, a monitoring framework for distributed systems that addresses both limitations by combining dynamic instrumentation with a novel relational operator: the happened-before join. Pivot Tracing gives users, at runtime, the ability to define arbitrary metrics at one point of the system, while being able to select, filter, and group by events meaningful at other parts of the system, even when crossing component or machine boundaries. Pivot Tracing does not correlate cross-component events using expensive global aggregations, nor does it perform offline analysis. Instead, Pivot Tracing directly correlates events as they happen by piggybacking metadata alongside requests as they execute. This gives Pivot Tracing low runtime overhead—less than 1% for many cross-component monitoring queries.
      pubtype: Periodical
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      dt:
        @attributes:
          year: 2020
    holdings:
      @attributes:
        islocal: N