Reliability is not something added after a system is built — it’s a set of decisions made during design about how the system behaves under load, how it fails, and how quickly a team can understand and recover when something goes wrong.
Traffic flows through the load balancer and application into the database, with monitoring watching the whole path — and feeding directly into incident response when something degrades.
Availability measures how much of the time a system is usable; reliability engineering is the discipline of deliberately designing, measuring and improving that number rather than treating it as an accident of how the system happened to be built. It applies engineering rigor — measurement, targets, trade-off analysis — to what used to be treated as purely operational firefighting.
A Service Level Indicator (SLI) is a specific measured metric — for example, the percentage of requests served successfully within 300ms. A Service Level Objective (SLO) is a target for that indicator, such as 99.9% over a rolling 30 days. The gap between 100% and the SLO is the error budget: the amount of unreliability the system is allowed before it's considered out of compliance. Error budgets are useful because they turn "should we prioritize reliability work or new features this sprint" into a data-informed decision instead of a debate.
Capacity planning means understanding how load is trending and provisioning ahead of it — enough headroom to handle growth and traffic spikes, without paying for permanently idle capacity. This depends on having accurate historical usage data and a reasonable growth forecast, not just reacting once a system starts approaching its limits.
Resilient systems assume components will fail and are designed to degrade gracefully rather than fail completely. Common patterns include timeouts and retries with backoff for calls to dependencies, circuit breakers that stop calling a failing dependency instead of piling up retries, and bulkheads that isolate failures in one part of a system from cascading into others.
When something does fail, a defined incident response process — clear ownership of who's leading the response, a communication channel, and a documented escalation path — reduces the time between detection and resolution. Post-incident reviews focused on what happened and what can be improved (not on blame) are what turn incidents into lasting improvements rather than repeated occurrences.
Disaster recovery planning defines how a system recovers from a large-scale failure — a full region outage, for example — including the recovery time and recovery point objectives discussed in infrastructure planning, and, critically, a tested procedure for actually executing that recovery rather than a document that has never been rehearsed.
Reliability and performance are closely linked: a system that's technically "up" but too slow to use is not meeting its users' actual needs. Monitoring needs to track both availability and performance against their respective targets, since either one degrading independently can represent a real reliability problem.
Operational readiness pulls these practices together into a single question worth asking before any system goes into production: if this fails at 3 a.m., does the team have the monitoring, the runbooks, the ownership and the tested recovery procedure to handle it — or would they be improvising for the first time under pressure?
A service level indicator (SLI) is a measured signal of user experience, such as the percentage of successful requests. A service level objective (SLO) is the target for that signal over a time window — for example, 99.9% of requests succeeding over 30 days.
The error budget is the amount of unreliability an SLO allows: with a 99.9% target, 0.1% of requests may fail within the window. Teams use it to balance the pace of releases against reliability work.
Clear ownership, meaningful health checks, observability that shows recent changes, safe and repeatable deployments, documented runbooks, and graceful degradation when a dependency fails.
Site Reliability
Reliability starts at architecture. Explore the practices that help teams build dependable systems and respond effectively when things fail.
More on this and related topics.
Production readiness is not a single checklist. It is the combination of reliability, security, observability, automation and operational discipline.
A practical guide to designing cloud infrastructure with the right balance of reliability, security, scalability and operational control.
A structured cloud migration starts with understanding applications, dependencies and infrastructure before moving workloads.