Pass The Foundation

Site Reliability Engineering (SRE) Explained (ITIL® 5)

Preparing for the ITIL® 5 Foundation exam? SRE is a term you'll see referenced alongside ITIL®, but it isn't an ITIL® practice — it's a distinct discipline that originated at Google. This guide explains what SRE actually is and how it connects to ITIL® 5.

Quick Answer

Site Reliability Engineering (SRE) is a discipline, born at Google, that applies software engineering approaches to IT operations problems, with the goal of improving the availability, performance, and reliability of production systems. SRE is not an ITIL® term or practice, but it overlaps conceptually with ITIL®'s Availability Management and reliability concepts, and it shares ITIL®'s emphasis on automation, though the two frameworks approach these goals from different traditions.

Where SRE Comes From

Site Reliability Engineering originated at Google, where it was described as what you get when you treat operations as a software problem. It's a job function, a mindset, and a set of engineering practices aimed at running reliable production systems. SRE teams are typically responsible for a combination of system availability, latency, performance, efficiency, change management, monitoring, incident response, and capacity planning. It is not an ITIL®-defined practice — it's a distinct discipline with its own origin, vocabulary, and practices.

What SRE Actually Focuses On

SRE brings software engineering skills to operational work, with heavy emphasis on automation and treating infrastructure as code. Rather than relying purely on manual operational processes, SRE teams write software to manage systems, fix problems at scale, and reduce repetitive operational toil.

How SRE Relates to ITIL® 5

SRE and ITIL® 5 aren't competitors — they address overlapping goals from different traditions. ITIL®'s Availability Management practice is concerned with ensuring services are available when needed, informed by concepts like reliability, maintainability, and serviceability. SRE pursues similar outcomes (system availability and reliability) but does so through an engineering-heavy, automation-first discipline built around specific practices like defining reliability targets and managing operational risk through engineering work rather than only through process governance. ITIL® 5 explicitly acknowledges DevOps-adjacent practices like SRE as relevant to how modern organizations actually run services, without redefining SRE as an ITIL® practice itself.

Real-World Example

A company running a large e-commerce platform has an SRE team responsible for the site's uptime and performance. That team writes automation to detect and remediate common failures, defines targets for how reliable the platform needs to be, and treats repeated manual firefighting as a problem to engineer away, not a permanent cost of doing business. The organization's separate ITIL®-aligned service management function still handles the broader Availability Management practice — setting availability requirements with the business and reviewing incidents that affected availability — with the SRE team's engineering work feeding directly into meeting those requirements.

Why This Matters

Understanding SRE matters because:

  • The exam expects you to recognize SRE as a distinct discipline, not an ITIL® practice
  • It illustrates how ITIL® 5 explicitly relates to and coexists with DevOps-adjacent disciplines, rather than ignoring them
  • It connects directly to ITIL®'s own reliability and availability concepts, giving you a concrete comparison point

Common Exam Mistakes

The most common mistake is treating SRE as an official ITIL® practice or assuming ITIL® invented the concept. SRE originated at Google, independently of ITIL®, as a distinct discipline.

A second mistake is assuming SRE and ITIL®'s Availability Management are the same thing described twice. They pursue related goals — reliable, available systems — but SRE is an engineering-heavy discipline with its own practices, while Availability Management is an ITIL® service management practice.

Memory Trick

Think:

SRE treats operations like software engineering.

ITIL® treats it like a governed service management practice.

Same destination — reliable, available services — different roads getting there.

Key Takeaways

  • Site Reliability Engineering (SRE) is a discipline that originated at Google, applying software engineering approaches to operations.
  • SRE is not an ITIL® term or practice — it's a distinct discipline with its own origin and vocabulary.
  • SRE and ITIL®'s Availability Management pursue related goals (reliable, available services) through different traditions.
  • SRE emphasizes automation and treating infrastructure as code to reduce manual operational toil.
  • ITIL® 5 explicitly acknowledges DevOps-adjacent disciplines like SRE as relevant to modern service delivery.

One Practice Question

Which statement best describes Site Reliability Engineering (SRE)?

  1. SRE is an official ITIL® 5 management practice.
  2. SRE is a discipline, originating at Google, that applies software engineering approaches to improve the reliability and availability of production systems.
  3. SRE and Availability Management are the exact same practice under two different names.
  4. SRE has no relevance to ITIL® 5 at all.
Show Answer

Correct Answer: B

SRE is a distinct discipline that originated at Google, applying software engineering practices to operations — it is not an official ITIL® practice, though it relates conceptually to ITIL®'s Availability Management and reliability concepts.

Frequently Asked Questions

Is SRE an official ITIL® 5 practice?

No. SRE originated at Google as a distinct discipline. ITIL® 5 acknowledges it as a relevant, related approach, but doesn't define or own it as an ITIL® practice.

How is SRE different from ITIL®'s Availability Management?

Availability Management is an ITIL® service management practice focused on ensuring services are available when needed. SRE is an engineering-heavy discipline that pursues similar reliability and availability goals through automation and treating operations as a software problem.

What does an SRE team actually do?

SRE teams focus on system availability, performance, latency, capacity planning, incident response, and reducing manual operational toil, often by writing automation and treating infrastructure as code.

Is this topic tested on the ITIL® 5 Foundation exam?

Yes, as part of understanding how ITIL® 5 relates to DevOps-adjacent disciplines like SRE.

Ready to Test Yourself?

Now that you understand what SRE is and how it relates to ITIL®, the next step is exploring ITIL®'s own concept of reliability. Take our free diagnostic quiz at PassTheFoundation.com to test yourself, or continue exploring the other ITIL® 5 core concept guides.