IEEE Transactions on Parallel and Distributed Systems · 1993 · 121 citations · 14 references
Cluster ComputingEngineeringVerificationConcurrent SystemEfficient ProtocolFault-tolerant MessagingFormal VerificationDependency RelationSynchronization ProtocolSystems EngineeringFault RecoveryConventional Synchronized ProtocolsParallel ComputingNetworked Computer SystemsDistributed SystemsComputer ScienceFault-tolerant NetworkDistributed ComputingScheduling (Operating Systems)Formal MethodsAsynchronous SystemsScheduling (Project Management)
The authors present an efficient synchronized checkpointing protocol that exploits the dependency relation between processes in distributed systems. In this protocol, a process takes a checkpoint when it knows that all processes on which it computationally depends took their checkpoints, hence the process need not always wait for the decision made by the checkpointing coordinator as in the conventional synchronized protocols. As a result, the checkpointing coordination time is substantially reduced and the possibility of total abort of the checkpointing coordination is reduced.< <ETX xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">></ETX>
14
K. Mani Chandy, Leslie Lamport · ACM Transactions on Computer Systems · 1985 · 2.4K citations · Full text
Engineering, Distributed Computing, Distributed Space Systems +12
Optimistic recovery in distributed systems
Rob Strom, Shaula Yemini · ACM Transactions on Computer Systems · 1985 · 722 citations · Full text
A message system supporting fault tolerance
Anita Borg, Jim Baumbach, Sam Glazer · ACM SIGOPS Operating Systems Review · 1983 · 245 citations · Full text
Engineering, Failover, Verification +18