Reliable Cron Across the Planet.
This article describes Google's implementation of a distributed Cron service, serving the vast majority of internal teams that need periodic scheduling of compute jobs. During its existence, we have learned many lessons on how to design and implement what might seem like a basic service. Here, we di...
| Published in: | Communications of the ACM Vol. 58; no. 6; pp. 48 - 54 |
|---|---|
| Main Authors: | , |
| Format: | Article |
| Published: |
Association for Computing Machinery
Jun2015
|
| Subjects: | |
| Online Access: | View this record in EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=102962468&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 102962468 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 00010782 ACM jtl: Communications of the ACM issn: 00010782 maglogo: N pubinfo: dt: Jun2015 vid: 58 iid: 6 pid: 68 pub: Association for Computing Machinery artinfo: ui: 102962468 10.1145/2732629 ppf: 48 ppct: 6 formats: tig: atl: Reliable Cron Across the Planet. aug: au: DAVIDOVIČ, ŠTĔPÁN GULIANI, KAVITA affil: Reliability engineer, Google on the Ads Serving SRE team charged with the reliability of the AdSense product Technical writer for technical infrastructure and site reliability engineers, Google Mountain View su: Utilities (Computer programs) Google Inc. Production scheduling Distributed computing Server farms (Computer network management) Reliability in engineering sug: subj: Utilities (Computer programs) Google Inc. Production scheduling Distributed computing Server farms (Computer network management) Reliability in engineering ab: This article describes Google's implementation of a distributed Cron service, serving the vast majority of internal teams that need periodic scheduling of compute jobs. During its existence, we have learned many lessons on how to design and implement what might seem like a basic service. Here, we discuss the problems that distributed Crons face and outline some potential solutions. Cron is a common Unix utility designed to periodically launch arbitrary jobs at user-defined times or intervals. We will first analyze the basic principles of Cron and its most common implementations and then review how an application such as Cron can work in a large, distributed environment, in order to increase the reliability of the system against single-machine failures. We describe a distributed Cron system that is deployed on a small number of machines, but is capable of launching Cron jobs on machines across an entire datacenter, in conjunction with a datacenter scheduling system. pubtype: Periodical doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y dt: @attributes: year: 2015 holdings: @attributes: islocal: N |
|---|