The Antifragile Organization.

The article discusses strategies for the prevention and management of software failures and system failures, particularly in large-scale distributed systems operating in cloud computing environments, arguing in favor of inducing system failures to develop system resilience during the design process....

Descripción completa

Detalles Bibliográficos
Publicado en:Communications of the ACM Vol. 56; no. 8; pp. 44 - 49
Autor principal: TSEITLIN, ARIEL
Formato: Artículo
Publicado: Association for Computing Machinery Aug2013
Materias:
Acceso en línea:Ver este registro en EBSCOhost
Descripción
Sumario:The article discusses strategies for the prevention and management of software failures and system failures, particularly in large-scale distributed systems operating in cloud computing environments, arguing in favor of inducing system failures to develop system resilience during the design process. Topics include redundancy, fault tolerance, Game Days, and autonomous agents known as "monkeys" that can be used to induce system failures. Companies mentioned include the video streaming company Netflix Inc.