SRE Fundamentals: SLOs, Error Budgets, and On-Call That Doesn't Burn People Out
You cannot buy reliability at the end of a sprint. It appears when your teams agree on user-centric targets, measure what matters, and use error budgets to steer day-to-day choices. Site Reliability Engineering gives you the language and practices to do that without trading developer velocity for fragile systems. This guide walks you from first principles to pragmatic implementation so you can run services that meet expectations, learn from incidents, and keep people healthy on call.
You will learn how to define SLIs and SLOs that reflect user experience, set and spend error budgets, build incident response muscle, and run postmortems that change behavior. You will see worked examples, math you can explain to product leaders, and templates that fit small teams as well as regulated enterprises. Along the way you will connect these practices to deployment, observability, capacity planning, and culture. If you need one place to align engineering, product, and operations on what reliability means and how to achieve it, use this page as your anchor.
What SRE is and why it matters
Site Reliability Engineering is an operations discipline rooted in software engineering. The premise is simple. Reliability is a feature, so you apply engineering rigor to deliver it just like you would for latency or cost. That means you specify reliability in terms a user would notice, like request success rate or page load time, then you automate measurement and enforcement. You also introduce safety constraints like error budgets that give teams a shared language to decide when to ship features and when to invest in stability. When done well, SRE reduces unplanned work and moves reliability from heroics to systems.
SRE sits inside the broader DevOps movement, but it is not a synonym for DevOps. DevOps focuses on breaking silos, fast feedback, and shared responsibility. SRE gives you a concrete set of practices, measurements, and guardrails to make that shared responsibility actionable. You can adopt DevOps values without SLIs or SLOs, but SRE asks you to write them down, automate the checks, and change your decisions when the numbers say so. If your organization is aligning on terminology and scope, it helps to place SRE within your existing DevOps strategy and to point newcomers at the DevOps learning hub for foundational context.
The other reason SRE matters is cost. Unreliability triggers churn, support tickets, and emergency work. It also creates decision paralysis, because teams ship fewer changes when they expect rollbacks and pager storms. SRE reduces this waste. When product leaders see reliability targets early, they plan features differently. When engineers see error budgets burning, they pause risky changes without debate. The system aligns around a limited resource, time without error, and uses that resource intentionally. You do not eliminate incidents, but you reduce their frequency and blast radius because you have clear definitions, fast detection, and consistent response.
Finally, SRE matters for people. On-call without SRE feels like being hunted by alerts. On-call with SRE feels like learning how to care for a system, with limits on toil, clear escalation, and off-ramps when load is too high. The difference is policy. You state which pages are valid, what you will not wake someone up for, and how you rotate, compensate, and learn. This guide will show you how to do that in a way that is humane, consistent, and supported by metrics.
SLIs, SLOs, and SLAs: definitions and relationships
Service Level Indicators are measurements of user-visible behavior. They answer questions like, what percentage of requests returned HTTP 2xx, or how long did it take for a customer to see the dashboard? You choose SLIs that reflect experiences users care about, not just low level counters. For a login API, a common SLI is successful_auth_responses / total_auth_attempts. For a dashboard, you might measure time to first meaningful paint for interactive page load. The key is to define these with a numerator and denominator so you can reason about error rates and budgets.
Service Level Objectives are targets for SLIs over a period. An SLO says, this SLI should be at or above a threshold for a given window. For example, 99.9 percent of POST /checkout requests will complete successfully in 30 days, and 99 percent of interactive dashboard loads will be faster than 2.5 seconds in a rolling 7-day window. SLOs are internal to engineering and product. They guide trade-offs but are not promises to customers. You should write SLOs in terms a product manager can understand and challenge, then publish them in your runbooks and dashboards.
Service Level Agreements are contracts with customers. They might include credits if reliability falls below an agreed threshold. SLAs tend to be looser than SLOs because you want breathing room. If your marketing site offers a 99.5 percent uptime SLA, you might operate at a 99.9 percent SLO so you are unlikely to breach the contract. Keep SLOs internal and more aggressive than SLAs, then set alerting on SLO burn rate, not on SLA failure. This avoids expensive surprises and keeps your teams focused on quality, not just legal minimums.
Here is a side by side comparison you can use to align stakeholders:
| Concept | Audience | Definition | Ownership | Typical format | Example |
|---|---|---|---|---|---|
| SLI | Engineering, product | A user-centric measurement of service behavior | Engineering | Ratio or percentile with scope | Successful requests / total requests, P99 latency |
| SLO | Engineering, product, leadership | A target for an SLI over time | Joint, with product | Threshold over a window | 99.9 percent success in 30 days |
| SLA | Customers, legal | A formal commitment with penalties or credits | Business, legal | Contract clause | 99.5 percent monthly uptime with service credits |
One more note on relationships. SLIs feed SLOs, and SLOs inform SLAs. The chain only works if measurement is sound. You cannot enforce a 99.95 percent SLO if you only sample at the wrong layer or exclude long tail users. Invest early in instrumentation and query definitions so that your SLO math matches how users feel. Later sections on instrumentation and query examples will give you a starting point you can adapt.
Choosing meaningful SLIs for your systems
Start with the user journey. Map what a user is trying to do, then list which interactions are core. For a SaaS product, typical journeys include sign up, login, view dashboard, create resource, update settings, and receive notifications. Pick the critical few, then define SLIs for availability and latency for each. When you can, measure from the user edge, not just the origin. A request that is fast in your data center but slow in a specific region because of a CDN policy is still a bad experience for those users.
For HTTP APIs, common availability SLIs include 2xx or 3xx responses divided by total requests. Decide how you treat 4xx. Many teams count 4xx as successful availability because the service responded with a valid status, even if the user sent a bad request. If you choose that, also define a separate SLI for functional correctness to capture spikes in 4xx that reflect regressions. For latency SLIs, measure percentiles, not averages. A median that looks fine can hide a P99 that is unacceptable. Good practice is to pick P90 for interactive UI loads and P99 or P99.9 for backend APIs where tail latency matters for batch jobs.
For asynchronous or streaming systems, availability is about throughput and lag. A data pipeline SLI might be records successfully processed divided by records ingested, with a separate SLI for end-to-end lag under a threshold, such as 95 percent of records are processed within 5 minutes of ingestion. Messaging systems benefit from SLIs tied to consumer backlog, for example, 99 percent of the time, lag is below 10,000 messages for topic X. Be careful to connect these back to user outcomes. A low backlog is only meaningful if downstream systems can keep up and users see fresh data.
For admin and maintenance operations, define SLIs for change success and time to recover. You might measure deploy success rate, percentage of rollouts without rollback, or time to restore service after a config change. These are user visible because failed deploys often cause incidents. You can also define SLI pairs to capture both success and speed, for example, 99 percent of schema migrations complete successfully, and 95 percent of migrations finish within 10 minutes. These meta SLIs support SLOs that improve your change process and reduce change failure rate over time.
When in doubt, write the SLI as a fraction. Good events divided by total events is durable. It supports error budget math and produces clean semantics for alerting. Here are a few examples:
- Web availability: SLI = count(status_code in [200, 399]) / count(all_responses) for GET /dashboard
- API latency: SLI = percentage of requests with latency < 300 ms at P99
- Data freshness: SLI = percentage of users whose dashboard reflects data no older than 15 minutes
- Deploy success: SLI = count(successful deploys) / count(total deploys) over 28 days
Designing SLOs with math you can explain
Choose an initial reliability target by combining business impact, competitive context, and engineering cost. A checkout API for an ecommerce site likely needs 99.95 percent monthly success, while an internal reporting UI might be fine at 99 percent. The cost curve is not linear. Each extra nine removes more error budget, which restricts change velocity and increases complexity. Be honest about the difference between wants and needs. You can always tighten later if users demand it and you have evidence that your change process can support it.
Define windows and thresholds before you pick alert policies. Fixed windows like calendar months are intuitive for leaders and finance, but they clip at month end. Rolling windows smooth this, for example 30-day sliding windows, but you must be explicit when reporting. A pragmatic approach is to define the SLO against a 30-day rolling window, then publish monthly and quarterly summaries for business stakeholders. For latency SLOs, combine a percentile and a threshold. For example, 99 percent of requests complete in under 500 ms over a 7-day rolling window.
Write SLOs in precise language. A complete SLO template includes scope, users included, requests counted, status code criteria, aggregation window, and threshold. Here is a template you can adapt:
- For the checkout service, for all authenticated users globally, 99.95 percent of POST /checkout requests will complete with HTTP status in [200, 399] over a rolling 30-day window.
- For the dashboard UI, 99 percent of interactive page loads will have time to interactive less than 2.5 seconds, measured at the browser using RUM, over a rolling 7-day window.
Quantify how SLOs compose. If two independent services each have 99.9 percent SLOs and a user flow requires both, the composite reliability is roughly 99.8 percent. This is why edge services that fan out should often have tighter SLOs, or you must add caching or retries so that the user journey is not limited by the product of reliabilities. When you design early, you can avoid architectural traps where a single slow dependency erodes your budget.
Finally, codify SLOs in version control. Treat them like product specs. Store definitions next to service code or in a centralized reliability repo. Use a schema or DSL so they are parsable, like OpenSLO style YAML. This lets you render dashboards, alert rules, and reports from one source of truth. It also supports peer review. Product and engineering can review changes to SLOs in pull requests, which builds shared ownership and prevents quiet downgrades to targets when things get hard.
Error budgets: from concept to daily decisions
An error budget is the amount of unreliability you are willing to tolerate in a period. If your SLO is 99.9 percent over 30 days, then your budget is 0.1 percent errors in that window. If you serve 1 billion requests, you can have up to 1 million errors. If you run 43,200 minutes in a 30-day month, you can be down for 43.2 minutes. Convert to the unit that makes trade-offs clear. For example, product leaders often reason better about minutes unavailable than percentages, so publish both.
Burn rate is how fast you are spending the budget compared to plan. If you have a 0.1 percent budget for 30 days, the planned daily spend is 0.1 percent divided by 30. If you are spending 0.05 percent per day, your burn rate is 0.5x and you are safe. If you have a 2 percent per day spike for 2 hours, your short window burn rate is 60x and you must act fast. Multi-window burn alerting uses two windows, a short one for fast spikes and a longer one for slow burns, to avoid noise. For example, alert when the 1-hour burn rate exceeds 14x and the 6-hour burn rate exceeds 6x. This catches both sharp incidents and creeping degradations.
Error budget policies turn numbers into behavior. Write down what happens when you have burned more than X percent of the budget in Y days. A simple policy ladder might be:
- At 25 percent budget burned early, require peer review on all deploys, plus canary on all changes.
- At 50 percent burned early, pause feature rollouts for 48 hours, focus on stabilization tasks.
- At 75 percent burned early, freeze changes except for critical fixes, run incident retro to plan mitigations.
- At 100 percent burned, maintain change freeze until budget resets or an executive accepts explicit risk.
These gates steer velocity without drama. Teams see the same numbers and know what will happen next. A flexible variation is a budget allocator. At sprint planning, set aside a portion of team capacity that will only be reclaimed if budget burn is under a threshold. Product leaders then have an explicit lever to trade feature work for resilience work.
Error budgets also shape alerting hygiene. You do not need a page for every failed health check if it does not affect the SLO. You do need a fast page for any condition that increases burn rate beyond your alert threshold. This conversion reduces alert fatigue. You collapse many micro-symptoms into a small set of budget-linked alerts and then build strong debugging playbooks to find root causes after the fact. We will cover instrumentation next so you can implement this with real queries.
Instrumentation and observability for SRE
You cannot operate against SLOs without trustworthy telemetry. At minimum you need service metrics for request outcomes and latency, logs to debug edge cases, and traces to see path segments across services. The RED method for services, rate, errors, and duration, is a simple frame. For each endpoint, expose counters for request rate, successful responses, and bucketed histograms for latency. Then build queries that compute SLIs over the chosen windows. If you are starting from near zero, prioritize the paths you picked as critical earlier and get usable metrics in place before automating everything.
Pick standard names and labels. A common pattern for Prometheus is to expose http_requests_total, labeled by method, path, status_code, and outcome, plus a histogram http_request_duration_seconds. For availability, derive success from status_code or a custom outcome label that marks application level success. For latency, derive percentiles from the histogram using quantile functions or use server side quantiles if provided by your time series system and library. If you are learning Prometheus from scratch, walk through the primer in Learn Prometheus without spending a cent to practice queries and recording rules.
Here are example PromQL snippets for an API availability SLI, a latency SLI, and a burn rate alert. Adapt label filters to your service names and paths.
# Availability SLI: ratio of successful to total requests over 5 minutes
sum(rate(http_requests_total{job="checkout",method="POST",path="/checkout",status=~"2..|3.."}[5m]))
/
sum(rate(http_requests_total{job="checkout",method="POST",path="/checkout"}[5m]))
# Latency SLI: percentage of requests under 300 ms at P99 over 5 minutes
histogram_quantile(0.99, sum by (le) (
rate(http_request_duration_seconds_bucket{job="checkout",method="POST",path="/checkout"}[5m])
)) < 0.3
# Error budget burn: current error rate divided by allowed budget rate, windowed
# Suppose SLO is 99.9 percent => budget = 0.001 errors per request over 30 days
# Compute over a 1 hour window for fast burns
(
1 - (
sum(rate(http_requests_total{job="checkout",path="/checkout",status=~"2..|3.."}[1h]))
/
sum(rate(http_requests_total{job="checkout",path="/checkout"}[1h]))
)
)
/
0.001
As you add services, push for consistent correlation across metrics, logs, and traces. Propagate a request ID or trace ID through your stack. Then, in an incident, you can pivot from a burn alert to a trace that shows which dependency is failing. A structured log with fields for request_id, user_id, and error_code lets you measure error budgets for specific cohorts or regions if needed. This is essential when your SLO is global but incidents are local, because you can decide to roll back a region while staying inside budget elsewhere.
Finally, connect observability to your SLO process explicitly. Build dashboards that show SLI time series, budget remaining, and burn rate at the top, with drill downs below. Link dashboards directly from your runbooks and readmes so on-call engineers land on the right views under stress. For a broader treatment of telemetry design, see the overview of observability for DevOps teams, then return here to wire those signals into SLO context and alert rules.
On-call that does not burn people out
Healthy on-call starts with scope and volume. If your service pages someone more than a few times a week during sleep hours, you have a design or policy problem. The first fix is to only page on symptoms that threaten the SLO. Everything else routes to tickets or email. This change reduces noise and reserves the pager for issues that require immediate human action. Tune alert thresholds based on burn rate or user-visible error rate, not hardware utilization. CPU at 95 percent might be fine under a burst, but a 20x burn for 30 minutes is not.
Design rotations to share load and protect recovery time. A common pattern is primary and secondary on-call, with a weekly rotation and no consecutive weeks for the same person. If the service is complex or high risk, add a shadow week before someone becomes primary, where the new engineer follows the pager and participates formally in incidents without the accountability. Document how to hand off at the end of a shift, including a brief note on any smoldering issues and action items. If you operate globally, consider follow-the-sun coverage to avoid waking people at night, but only if you can hand off cleanly between regions and teams.
Runbooks are your on-call safety net. For each SLO alert, write a runbook that includes a one line summary, how to validate the alert, likely causes, step by step diagnostic commands or dashboards, rollback or mitigation actions, and escalation criteria. Keep them short and direct. A good test is whether a capable engineer from a neighboring team can follow the runbook under pressure without paging a subject matter expert immediately. Store runbooks with your service code and version them alongside changes. Link them from the alert so the on-call can click straight to the right instructions.
Compensation and limits matter. Pay for on-call. Set a maximum page volume per person per week. If someone exceeds that limit, remove them from rotation and compensate for the overload. Rotate incident commander duty separately from service ownership to keep leadership muscle distributed and to give people breaks from decision fatigue. Encourage people to page early when unsure. Punish people for not paging only when the situation warranted it and the playbook called for it. The goal is a culture where on-call engineers act with confidence within well defined guardrails.
Finally, measure on-call health. Track pages per week, time spent on incident response, after hours interrupts, and follow-through on postmortem action items. Review these metrics in team retrospectives. Use them to trigger toil reduction work, staffing changes, or rotation adjustments. A humane on-call is not an accident. It is an outcome of instrumentation, policy, and leadership choices. Write those choices down and revisit them as your service evolves.
Incident response and incident command
Incidents are unplanned events that disrupt your service level objectives. Treat them like fire drills with a playbook. Your incident response process should define severity levels, roles, communications, and timelines. When an alert fires that meets your incident criteria, you open an incident channel, assign roles, and begin triage. Things move fast, so keep roles simple and their responsibilities clear.
A minimal role set includes incident commander, operations lead, communications lead, and subject matter experts. The incident commander runs the process, keeps a timeline, and makes decisions. The operations lead drives technical work, assigns tasks, and coordinates mitigation. The communications lead posts regular updates to stakeholders, both internal and external if needed. Subject matter experts diagnose, test hypotheses, and implement fixes. In small teams, one person may play several roles, but remember to split IC and ops lead if cognitive load gets high.
Define severities and expected response. Keep it simple and map to user impact and SLO risk. A typical matrix looks like this:
| Severity | User impact | Required response time | Roles engaged | Communication |
|---|---|---|---|---|
| SEV-1 | Broad outage, critical flow blocked, SLO at extreme risk | 5 minutes to acknowledge, continuous engagement | IC, ops lead, SMEs, comms lead | Exec updates every 15 minutes, external status if needed |
| SEV-2 | Partial outage or major degradation, SLO at risk | 15 minutes to acknowledge | IC, ops lead, SMEs | Stakeholder updates every 30 minutes |
| SEV-3 | Minor degradation or localized impact, SLO safe but trending | 1 hour to acknowledge | Ops lead, SMEs | Team updates at resolution |
Codify incident stages. A simple flow is detect, declare, mitigate, resolve, follow up. Detect is the alert or report. Declare is when you open the incident and assign a severity. Mitigate is when you take the fastest safe action to restore service, like rollback, failover, or feature flag disable. Resolve is when user impact is back to normal and SLOs are no longer at risk. Follow up is the postmortem and action item assignment. Write timelines during the incident so you do not rely on memory later. Many chat tools support slash commands to stamp events. Use them.
Communicate proactively. For SEV-1 and SEV-2, post updates on a regular cadence even if the update is that you are still investigating. This prevents leader and customer panic, and it gives space to engineers to work. When you share externally, be honest and avoid speculation. Acknowledge impact, share knowns, state next update time, and apologize. Internally, link to dashboards that show SLO impact so product and support see the same view. This builds trust and keeps everyone aligned on the goal, which is to restore service and avoid further harm.
Postmortems that lead to learning
A postmortem is a written analysis of an incident that explains what happened, why it happened, how you restored service, and what you will do to reduce risk. Blamelessness is a requirement, not a preference. You will only get honest narratives and real fixes if people know they will not be punished for making reasonable mistakes in complex systems. Focus on system conditions, decision points, and feedback loops. Ask, what made the error possible and what made it likely? Then propose changes that alter those conditions.
Use a consistent template. Useful sections include summary, impact measured in SLI terms and user stories, timeline, detection and response analysis, contributing factors, what went well, what did not, and action items with owners and due dates. Include graphs from the incident that show SLO breach and recovery. Include links to relevant runbooks and code changes. Keep the document short enough to read in 10 minutes, but detailed enough that someone could learn from it a year later without extra context. Review drafts in a meeting that includes engineering, product, and support.
Classify action items. You will have quick wins, like fix a missing alert or update a runbook, and deeper investments, like build a replayable staging environment or add graceful degradation to a dependency. Tag items by category, for example, detection, mitigation, prevention, or resilience. Estimate effort and expected risk reduction. Then place them into your reliability backlog. Link these items to your error budget policy so that when burn rate is high, you have a prepared queue of valuable work to pull.
Close the loop. Track postmortem action item completion rates and tie them to incident frequency over time. When you see recurring themes, invest in root cause classes. If multiple incidents share configuration drift as a contributor, prioritize infrastructure as code and policy-as-code. If several point to risky manual rollouts, adopt stronger CI and CD practices or GitOps workflows. Publish a quarterly review to leadership that shows incident trends, SLO conformance, and completed improvements. This converts incident pain into organization-level learning.
Toil reduction and automation
Toil is manual, repetitive, automatable work that grows with service size. Toil crowds out engineering time for improvements. Track toil explicitly. Ask engineers to label tasks as toil or engineering in their weekly summaries. Measure hours spent on paging, manual deploys, ticket processing, and ad-hoc data fixes. Then set a goal to reduce toil percentage over time. Teams that consistently invest 20 to 30 percent of their time in toil reduction compound gains, because each automation frees time to automate the next item.
Start with high frequency, low risk tasks. Automated rollbacks, log triage scripts, and routine maintenance can often be scripted in a day. Use your incident timeline to find the most common copy-paste commands. Wrap them in a safe script with guardrails, then document usage in the runbook. For infrastructure, adopt infrastructure as code so that environment changes are reviewed, versioned, and reproducible. Combine IaC with policy-as-code to prevent drift and to catch risky configurations before they reach production.
Bring automation into your release pipeline. Manual deployments are classic toil, and they often contribute to change failure rate. Set up continuous integration to run tests and build artifacts on each change. Set up continuous delivery with stages for canary and automated health checks. These stages should consult your SLOs or proxy health metrics to decide whether to advance. A well tuned CI and CD pipeline reduces human decision points and shortens mean time to restore because rollbacks are one click or automated when health checks fail.
Use Git as the interface for operations where possible. GitOps practices sync desired state declared in repositories to clusters and services automatically. This eliminates a class of manual config pushes and gives you a full audit trail for changes. Combine with chatops to run common actions from your incident channel with fine grained permissions and templates. The pattern is the same. Move repetitive steps into code, add review and tests, and let systems do what they do best, which is consistent execution.
Finally, do not confuse toil reduction with zero manual work. You still need humans for judgment and for rare creative handling. The goal is to remove manual work that does not need that judgment so you increase the ratio of interesting problems to repetitive tasks. As you reduce toil, on-call gets quieter, incidents become easier to handle, and you create space to work on deeper reliability features like graceful degradation or adaptive concurrency controls.
Reliability in the SDLC and release engineering
Reliability is shaped by how you build and release software. Small, frequent changes are easier to reason about, test, and roll back. Adopt feature flags so you can ship code dark and activate functionality gradually. Use flags to decouple deploy from release, which lets you test in production gently and reduce blast radius. Integrate flags with your SLOs by pausing progressive rollouts when burn rate increases. Make this automatic where feasible so humans do not have to make high stakes calls under pressure.
Progressive delivery techniques like canary and blue-green deployments reduce risk. In a canary, you send a small percentage of traffic to a new version and compare key SLIs, such as error rate and latency, against the baseline. If the canary looks good, you increase the percentage. If it looks bad, you roll back. This is much easier in orchestrated environments. If you are running containers, review Kubernetes fundamentals for DevOps teams to understand how Deployment strategies and Horizontal Pod Autoscalers affect reliability. Your release engineering should know how to roll out without exhausting resources or creating cascading retries.
Change control should be lightweight but visible. Require peer review with checklists that include reliability questions. What is the blast radius, are there runbook updates, is there an automated rollback? For risky changes, do load testing in an environment that mirrors production enough to catch obvious regressions. If you work in a regulated space, define a traceable change management record but automate as much of it as possible. Use your CI and CD system to capture build provenance and approvals so auditors do not create parallel processes that slow engineering without improving safety.
Connect SLOs to release decisions with policy. If budget burn is above your alert thresholds, slow or pause changes. If you must ship urgently, require executive approval with an explicit risk statement. This approach places SRE and product in a partnership. Product understands that reliability is a feature that keeps customers, and SRE understands that shipping value matters. The error budget is the contract between these goals. When you respect it, you can move fast most of the time and slow down when needed without politics.
Capacity planning and performance engineering
Reliability depends on headroom. If your system runs too close to saturation, small traffic bursts or noisy neighbors will push you into failure modes. Define SLIs for saturation, like 95th percentile CPU usage below 80 percent, queue depth below a threshold, or garbage collection time under a limit. These are not user-visible by themselves, but they predict user-visible errors. Connect these to autoscaling where appropriate. For services on Kubernetes, tune resource requests and limits so the scheduler has room to place pods, and configure autoscalers based on request rate or custom metrics, not just CPU.
Load test your critical flows regularly. Use production-like data and realistic traffic patterns, including burstiness and region skew. Measure SLI degradation as load increases, not just raw throughput. Your goal is to find the knee of the curve where latency or error rate accelerates. Then set thresholds and alerts far enough from that knee to absorb common bursts. If you rely on caches, include cache warm up and eviction scenarios. Many outages hide in the transitions when a cache rebuilds or a warm path goes cold after a deploy.
Architect for graceful degradation. When a dependency is down or slow, prefer to shed non-critical load, serve cached or stale data, or disable secondary features rather than failing entirely. Design bulkheads between features so one slow service does not block unrelated requests. Tune timeouts and retries carefully. Too aggressive retries under partial failures will amplify load and cause cascades. A simple rule is to cap retries with jitter, use exponential backoff, and honor idempotency. Document these patterns in your service templates so they are consistent across teams.
Capacity planning is not a yearly meeting. It is a rolling practice where you compare current trends with your growth forecasts and your SLO targets. Publish a monthly capacity report with headroom, projected saturation dates, and planned mitigations like hardware purchases or architecture changes. Connect this to product roadmaps. If marketing plans a campaign, coordinate to ensure headroom exists or to prepare circuit breakers that protect core flows when traffic exceeds safe limits. Reliability improves when these conversations happen ahead of time instead of in the middle of a promotion.
SRE at different scales
At startup scale, keep SRE practices lightweight and pragmatic. Start with one or two critical SLIs per key flow, a single SLO per service, and an on-call rotation that includes the engineers who build the system. Use error budgets to pause risky changes when necessary, but accept that some messiness is part of learning. Focus on instrumenting core paths, automating deploys, and writing runbooks for the top alerts. Invest early in a small set of safety features like feature flags and rollbacks. Over time, layer on incident command formality as page volume grows.
At mid-size, separate roles and formalize process. Create an SRE team that partners with product teams to define SLOs and build platform capabilities. Keep ownership with the service teams, but provide tools, templates, and coaching from SRE. Establish clear incident command, a regular postmortem process, and an error budget policy reviewed by engineering leadership. Build cross-team dashboards for organizational SLO status so leaders can see where to invest. Balance shared infrastructure like observability and CI with team autonomy.
In regulated or enterprise environments, you will have additional constraints around change management, audit, and incident reporting. Integrate these requirements into your SRE process so they add value instead of friction. For example, generate change tickets automatically from your CI and CD pipeline, record approvals in version control, and use SLO dashboards as part of compliance evidence. Train incident commanders to handle external regulators or customer communications in partnership with legal. Invest in chaos engineering and disaster recovery drills to satisfy recovery objectives while improving real resilience.
Platform SRE emerges as you scale to dozens of teams. The platform team provides paved roads for deployment, telemetry, and security. Their SLOs are internal and support the developer experience, such as build times or deployment success. Service SRE works with individual teams on their user-facing SLOs. The two groups share tooling and standards but have different customers. Align them through common libraries, GitOps workflows, and playbooks so that when a product team builds a new service, they inherit healthy defaults rather than reinventing every decision.
Building the SRE practice: staffing, skills, and career paths
Staff SRE by mixing software engineering depth with operations sensibility. Look for people who can code, automate, and reason about systems under stress. At junior levels, emphasize debugging skills, willingness to learn, and healthy on-call habits. At senior levels, add design skills for resilience patterns, the ability to run large incidents, and influence across teams. Avoid treating SRE as a pure ops role. SREs should spend significant time on engineering improvements, not just running the pager.
Train continuously. New SREs need hands-on practice with your stack and incident process. Pair them with experienced on-call engineers. Run game days that simulate failures so they can practice command and technical response without real user impact. Encourage certifications only if they map to skills you need. If you are building a DevOps and reliability foundation, see the overview of industry recognized DevOps certifications to plan learning paths that complement on-the-job experience. Pair that with readings, internal design reviews, and retrospective facilitation practice.
Create a growth path that values both depth and breadth. SREs can become staff engineers who design organization-wide reliability patterns, or managers who lead the SRE function. They can also rotate into product teams to seed SRE habits, then return. Treat this mobility as a strength. It spreads reliability thinking into development and brings product empathy back into SRE. Support non-traditional backgrounds. Operations analysts, QA engineers, and developers from adjacent domains can become strong SREs with targeted mentoring and practice.
If you want structured preparation, consider formal learning plus applied project time. Many engineers accelerate by joining programs that combine study, labs, and real-world sprints. If that fits your current goals, explore the DevOps Engineer study and internship program that includes hands-on SRE practice. You will get guided exposure to SLIs, SLOs, pipelines, infrastructure as code, and incident response in a format that complements daily work. Pair such programs with internal mentorship so learning transfers to your stack.
Finally, keep leadership aligned on reliability trends and talent needs. Share hiring pipelines, on-call health metrics, and postmortem themes in quarterly reviews. Use research and market signals to inform your roadmap. If you are scanning the horizon, the analysis in the DevOps trends, tools, and career guide can help you anticipate which practices, tools, and skills will matter most in the next planning cycles.
Putting it together: a worked example SLO program
Imagine you run a multi-tenant SaaS with two critical user flows, login and dashboard view, and one critical admin flow, deploy. You choose SLIs as follows. For login, availability is successful_auths / total_auths with success defined as HTTP 2xx and valid token issued. For dashboard, latency is time to interactive measured by a RUM script injected into the page. For deploy, success is successful_deploys / total_deploys. You define SLOs as 99.95 percent monthly for login availability, 99 percent of dashboard loads faster than 2.5 seconds over 7 days, and 99 percent deploy success over 30 days.
You turn SLOs into error budgets. For login, 0.05 percent monthly errors on 200 million requests allow 100,000 bad attempts, whether due to backend errors or third-party identity issues. For dashboard, 1 percent of loads can exceed 2.5 seconds over a 7-day window. For deploy, 1 percent can fail, but you prefer to use this as a leading indicator for toil and to reduce change failure rate. You publish these budgets to product and set an error budget policy that pauses new feature flags when login burn exceeds 50 percent in the first 10 days of a month.
You implement instrumentation. Your API emits http_requests_total and http_request_duration_seconds with labels for method, path, and outcome. Your web app publishes a RUM beacon with time to interactive for the dashboard, including a user_id and region label. Your CI publishes deploy events with a success flag to a metrics sink. You write PromQL recording rules for availability, latency percentiles, and deploy success by service and environment. You build a dashboard that shows SLO compliance, budget remaining, and burn rates for 1 hour and 6 hour windows.
Alerts fire when burn breaches thresholds. One morning the 1-hour login burn rate hits 16x and the 6-hour hits 8x. The alert pages the primary. They open an incident, assign roles, and check the dashboard. The availability SLI shows a spike in 5xx for POST /auth. Traces show increased latency in a dependency call to a third-party identity provider. Runbook steps suggest enabling a local cache of keys and reducing timeout to avoid cascading retries. The team implements the mitigation and burn falls below 1x within 30 minutes. Later, the postmortem identifies the missing timeout and the lack of circuit breaker as contributors. Action items include adding a bulkhead for the identity provider and improving synthetic checks against the provider to catch slowness earlier.
Meanwhile, deploy success SLI shows increased failures on Friday evenings. This is not paging, but the postmortem trends suggest that deploys close to end of day have higher failure rates and cost weekend on-call time. The team updates the deploy policy to restrict deploys after 4 pm local time unless an on-call engineer volunteers and a rollback window exists. They also add a canary step for schema migrations with read-only shadow tables. The SLO for deploy success improves, and incident frequency drops because fewer risky changes happen when support is thin.
Over a quarter, you review SLO performance. Login SLO was met at 99.951 percent, dashboard latency SLO was missed twice following large frontend framework updates, and deploy success was 99.4 percent. Product recognizes that the dashboard misses correlate with big UI changes and approves a capacity increase and prefetch strategy during rollouts. SRE reports that on-call pages dropped by 35 percent and average time to resolve fell by 20 percent due to better runbooks and automation. The error budget process became normal. Engineers paused a rollout in week two without argument, because the budget numbers and policy were clear.
Practical math for SLOs and budgets
Make SLO math visible and simple. For availability SLO A over period T, with total events N, error budget E equals (1 - A) * N. If your events per second vary, estimate N from historical data or model it as a range. For time-based SLOs like uptime, where the SLI is minutes available, the error budget in minutes M is (1 - A) * minutes_in_T. For a 30-day period, minutes_in_T is 30 * 24 * 60 = 43,200. For A = 0.999, M = 43.2 minutes. Use these numbers in planning so leaders grasp the cost of adding another nine. Moving from 99.9 percent to 99.99 percent cuts budget from 43.2 minutes to 4.32 minutes. That demands redundancy and slower change.
Use multi-window, multi-burn alerts to keep signal high. Define budget rate per minute as budget_per_minute = (1 - A) / total_minutes_in_T. Compute observed error rate over window W as observed_error_rate_W. Burn rate BR_W = observed_error_rate_W / budget_per_minute. Alert when BR_1h > 14 and BR_6h > 6. These numbers are common starting points, because they catch about 2 percent budget spend in an hour or 5 percent over 6 hours. Adjust for your patterns. If you get false positives during planned maintenance, refine your SLI scope or label maintenance windows and mute alerts.
For latency SLOs, modeling takes care. If your SLO is P99 less than 300 ms, the error budget is the area where latency exceeds 300 ms under the P99 curve. You can approximate errors as count of requests above threshold in each bucket divided by total, but be careful with histogram bias. When possible, export explicit counts for threshold breaches. If not, sample percentiles with caution and validate against synthetic checks or logs. Explain these subtleties to product once so they understand that latency SLOs have more measurement error than plain availability.
Composite SLOs are tricky. If your user journey needs two services A and B, and both must succeed, composite reliability is roughly A * B if independent. If they share dependencies, the composite is worse than that product. To manage this, set tighter SLOs for the most central services or build client-side fallbacks. For example, if the recommendations service is slow, degrade the home page to render without recommendations within 500 ms. This keeps the page within its latency SLO while the recommendations team fixes its issue. Document such interactions and revisit them in architecture reviews.
Finally, treat SLOs like financial plans. Maintain a ledger of budget spend with annotations for big draws, like SEV-1 incidents or high traffic campaigns. At the end of the period, publish a report with budget left, budget spent by cause, and improvements delivered. Over time, this helps justify investments in resilience. Leaders see that a one-time spend on a circuit breaker paid back in budget saved and in feature velocity maintained.
Service catalogs, ownership, and SLO governance
SLOs only work when you know who owns them. Maintain a service catalog that lists each service, its owners, its critical paths, its SLOs, and its runbooks. Include links to dashboards and code. Keep the catalog lightweight but current. Automate updates from code repositories and on-call rotas where possible. Make it easy for product, support, and incident commanders to find the right page to ping.
Create an SLO review cadence. Each team reviews their SLO performance quarterly with product stakeholders. They propose adjustments to targets based on user feedback, competitive shifts, or engineering improvements. They retire SLOs that no longer matter and add new ones for features that now drive user value. A reliability steering group reviews cross-cutting risks and allocates shared investment, like platform work to improve observability or deploy speed. Keep meetings short and data-driven. Show SLI histories, budget burns, and incident impacts.
Publish SLOs in places where they are used. Put them in the team readme, in the runbook, and on the top of the service dashboard. Link to them from story templates so that when a change lands, the reviewer asks how it affects SLOs. Train product managers to set goals that reference SLOs. For example, grow traffic by 20 percent while preserving dashboard latency SLO adherence at 99 percent. This reframes success as value plus quality, not value alone.
Finally, avoid SLO sprawl. Too many SLOs dilute focus and confuse alerting. Start with one or two per critical flow and add only when you see a distinct class of incidents that a new SLO would help prevent. Evaluate SLOs annually. Remove the ones that no longer drive behavior change. Keep the ones that teams consult when making decisions. This keeps your SRE practice simple and effective.
Tools, templates, and representations
You can represent SLOs in code to generate alerts and reports. A common approach is a small YAML schema that captures the SLI type, the target, the window, and the alert policy. A simplified example looks like this:
service: checkout
slo:
name: post-checkout-availability
description: 99.95 percent of POST /checkout succeed in 30 days
sli:
type: ratio
numerator: sum(rate(http_requests_total{job="checkout",path="/checkout",status=~"2..|3.."}[5m]))
denominator: sum(rate(http_requests_total{job="checkout",path="/checkout"}[5m]))
objective: 99.95
window: 30d
alerts:
- short_window: 1h
long_window: 6h
short_burn_threshold: 14
long_burn_threshold: 6
From this, generate Prometheus recording rules, Grafana dashboards, and alert manager entries. Keep the generation step in your CI so changes to SLOs go through review and build. This reduces drift between definitions and operational reality. If you operate across teams, publish a template repo that new services can fork, with stubs for SLIs, alerts, runbooks, and incident roles.
Maintain runbook templates as well. A minimal template covers alert context, validation, dashboards, common causes, commands, rollbacks, and escalation. Include a section for safety checks, like verify you are in the correct environment, and list commands that can cause additional harm. Add a section for post-incident cleanup, like restore autoscaling, re-enable feature flags, or remove temporary firewall rules. Keep the tone direct and specific. Under stress, engineers should not have to interpret vague guidance.
Finally, build a small library of incident macros. Pre-filled chat messages that declare severities, assign roles, and post update cadence save seconds that matter. A macro like /incident declare SEV-2 Login degradation, 5xx elevated for POST /auth in region us-east, IC @alice, Ops @bob, Comms @charlie, next update in 15 minutes creates order quickly. Combine with chat integrations that post SLO graphs and burn rates when you declare. You do not need a big platform to improve incident response. A few well placed automations make a big difference.
Connecting SRE to platform and architecture choices
SRE is not a layer you bolt on after architecture is set. It should inform platform and design choices from the start. When you choose a database, evaluate not just throughput but backup and restore behavior, failover modes, and metrics emitted. When you choose a message queue, test what happens under consumer backpressure and how dead letter queues behave. Ask vendors for their own SLOs and hold them to error budgets. Record vendor incidents in your postmortems and model composite SLOs that include third-party dependencies.
Infrastructure platforms shape incident response. If you run on Kubernetes, you inherit scheduler behavior, readiness and liveness probe semantics, and controller loops that can help or hurt during failures. Learn these patterns so you do not fight them. If you host on cloud managed services, understand regional isolation, maintenance windows, and quotas. Build playbooks for provider outages that include traffic shifting and degraded modes. Connect architectural durability, such as multi-region replication, to user-facing SLOs. Do not pay for multi-region if your SLOs and business priorities do not require it, but do not avoid it if your composite reliability demands it.
Build reliability patterns into service templates. Include client libraries that implement timeouts, retries, and circuit breakers with sane defaults. Include structured logging, tracing, and metrics instrumentation out of the box. Provide easy feature flagging and rollout scripts. Make the reliable path the default so product teams do not have to rediscover the same techniques. The best SRE outcomes happen when teams rarely need to think about plumbing and can focus on business logic and user needs.
Finally, measure the platform itself. Your CI and CD system has SLOs, like build time under 10 minutes for 95 percent of jobs and deployment success at 99 percent. Your observability stack has SLOs, like metrics scrape success and dashboard load time. Publish these internally. When platform SLOs slip, it affects everything else. It also burns goodwill. Investing in platform reliability pays back as reduced friction, faster incident detection, and better morale across product teams.
FAQ: SRE fundamentals
Q: What is the difference between SLI, SLO, and SLA in one sentence each? A: An SLI is a measurement of user-visible service behavior, an SLO is a target for that measurement over time used by engineering and product, and an SLA is a legal commitment to customers that usually sits below your internal SLO.
Q: How many SLOs should a team start with for a new service? A: Start with one availability SLO and one latency SLO for the most critical user flow, then add more only when they drive different decisions. Too many SLOs dilute focus and create alert noise.
Q: Do I need to page for every SLO violation? A: Page for conditions that threaten to burn your error budget quickly, not for every small deviation. Use multi-window burn rate alerts to catch fast spikes and slow leaks without waking people unnecessarily.
Q: How do I pick a reliability target like 99.9 percent versus 99.99 percent? A: Choose based on user impact, competitive expectations, and engineering cost. Show leaders the error budget math in minutes or allowed errors, then align on how much change velocity you are willing to trade for the extra nine.
Q: Should 4xx responses count as errors in availability SLIs? A: Many teams count 4xx as successful availability because the service responded correctly to a bad request, but they track 4xx trends separately for functional regressions. Document your choice and be consistent across services.
Q: How do I keep on-call humane while services grow? A: Only page on SLO threats, design fair rotations with backup and handoff, invest in runbooks and automation, measure page volume and after hours interrupts, and set explicit compensation and overload limits.
Q: What does a blameless postmortem look like in practice? A: It documents the incident clearly, focuses on system conditions and decision points rather than individual mistakes, includes impact in SLI terms, and ends with prioritized action items that reduce risk. It is shared widely and used to adjust process and architecture.
Q: How do SRE and DevOps relate in an organization? A: DevOps provides cultural principles like shared responsibility and fast feedback, and SRE provides concrete practices like SLIs, SLOs, error budgets, incident command, toil reduction, and automation to make those principles operational. Point teams at the DevOps hub for context and use SRE to implement the habits that produce reliability.
