Chapter 13
Observability, Failure & Recovery
Logs, metrics, and traces to see what a system is doing, plus failure-recovery patterns — retries, circuit breakers, timeouts — that contain bad services.
What You'll Learn in This Chapter:
- ✦Heartbeats and Gossip protocols
- ✦Circuit Breakers and Retry policies
- ✦Distributed tracing and metrics collection
- ✦Graceful degradation strategies