Uncategorized

CloudOps, DevOps, SRE and Platform Engineering: Who Owns What?

Close-up of rack-mounted server equipment in a Wikimedia Foundation server rack.

Quick answer

Divide operational work by service, decision and action, not by job title alone. Name an accountable owner, an execution team, an approval authority where needed, and an escalation route. Application teams can own their workloads while CloudOps, SRE and platform teams provide clearly bounded operational capabilities. Record those boundaries in a discoverable agreement, including who resolves ownership gaps. AWS operational guidance explicitly recommends documented responsibilities and escalation to someone authorized to assign ownership. (docs.aws.amazon.com)

1. What the four disciplines mean in practice

These terms describe overlapping approaches to delivering and operating software. They are not four mutually exclusive departments.

CloudOps is used here as a practical label for operating cloud environments: maintaining operational visibility, managing changes, supporting infrastructure and coordinating operational improvements. Its precise remit should be agreed locally. AWS describes CloudOps through a business-aligned operating model, supported by observability, safe automation and reversible changes—not through a mandatory organization chart. (docs.aws.amazon.com)

DevOps emphasizes collaboration across development and operations throughout the software lifecycle. It is broader than maintaining a deployment pipeline. Google’s SRE literature describes DevOps as principles for whole-lifecycle collaboration, with culture, measurement and change management playing important roles. (sre.google)

Site reliability engineering, or SRE, applies software engineering to operational problems. Google’s formulation emphasizes engineering systems that replace manual operational work. Its practices include managing services against service-level objectives, or SLOs: agreed reliability targets appropriate to the service and its users. (sre.google)

Platform engineering plans and provides common computing capabilities for internal users, including application teams. CNCF’s definition includes people, processes, policies and technology—not just a developer portal. In AWS’s cloud operations and platform enablement model, platform engineers turn standardized patterns into self-service capabilities. (tag-app-delivery.cncf.io)

A useful working distinction is:

Discipline Question it helps answer
CloudOps How do we operate this cloud environment effectively?
DevOps How do development and operations collaborate to deliver changes?
SRE How do we engineer and manage the required service reliability?
Platform engineering Which reusable capabilities help application teams work independently?

Use these questions as planning prompts, not rigid boundaries. One engineer may contribute to all four disciplines.

2. Separate accountability, execution and approval

“Who owns production?” is too broad to produce a useful answer. Break it into narrower questions:

  • Who decides whether an application release is ready?
  • Who maintains the deployment mechanism?
  • Who can approve production access?
  • Who responds when the application fails?
  • Who investigates unexpected spending?
  • Who accepts an unresolved operational risk?

AWS distinguishes several meanings of ownership, including control over changes, troubleshooting support, and financial or administrative accountability. It recommends recording owners for workloads, infrastructure, platforms and applications. (docs.aws.amazon.com)

For your ownership agreement, use four explicit fields:

Field Meaning in this agreement
Accountable owner Ensures the defined outcome is addressed and unresolved work is tracked
Execution team Performs the deployment, investigation, approval processing or repair
Approval authority Authorizes a restricted action or exception under local policy
Escalation route Receives unanswered requests, disputed boundaries or missing ownership

The accountable owner need not perform every task. Equally, the team running an automation does not automatically own the business decision behind it.

Avoid a single row called “deployments.” Separate application releases from shared delivery infrastructure. Similarly, separate access policy, individual access approval and technical provisioning.

Treat this as an operational worksheet rather than a universal framework. AWS specifically recommends responsibility matrices and a defined mechanism for resolving ownership gaps. (docs.aws.amazon.com)

3. An ownership matrix for a small application team

The following matrix is illustrative, not a prescribed staffing model. It assumes one application team, a technical lead, an operations engineer and a designated budget owner. One person may hold several roles.

Activity Accountable owner Execution Approval or escalation
Application release and rollback Application lead Application engineers using the agreed workflow Apply release policy; escalate uncertain recovery to the technical lead
Deployment workflow and infrastructure automation Technical lead Operations engineer with application contributors Review under the local change policy
Production access request Designated resource owner Authorized access administrator Policy-designated approver; requester does not self-approve in this example
Application incident Application lead for follow-up Assigned responder, supported by application and operations engineers Escalate to a named backup and incident lead
Cloud infrastructure incident Technical lead Operations engineer with affected application engineers Escalate to the backup responder and provider support when appropriate
Unexpected cloud spending Designated budget owner Operations engineer investigates usage; application engineers investigate workload behavior Budget owner decides business trade-offs
Backup procedure and restoration check Application lead Operations engineer and application engineer Escalate an unverified recovery path to the technical lead

This arrangement deliberately avoids creating four separate teams. AWS’s enablement model allows operational expertise to come from a dedicated team or a virtual team, with application teams progressively taking responsibility for their systems. (docs.aws.amazon.com)

Before adopting the matrix, check whether each assignment is executable:

  • Does the responder have the necessary permissions and tools?
  • Can another suitably skilled person follow the procedure?
  • Is there an available backup?
  • Does the documented support coverage match actual staffing?

AWS’s procedure-ownership guidance specifically asks owners to verify that an adequately skilled team member can perform an activity with the correct access, permissions and tools. (docs.aws.amazon.com)

If coverage is limited, document that limit plainly. A matrix is not evidence of round-the-clock support.

4. An ownership matrix for a larger platform organization

The second illustrative model separates application ownership from shared capabilities. It assumes application teams, a platform team, CloudOps, an SRE engagement, an access-governance function and budget owners.

Activity Accountable owner Execution Approval or escalation
Application release and rollback Application service owner Application team through supported delivery capabilities Release policy; escalate workflow failures to platform
Shared delivery service Platform service owner Platform team Shared-service change policy; escalate impact to incident coordination
Cloud account and shared infrastructure operations CloudOps service owner CloudOps team Relevant change authority; governance exceptions follow a separate route
Production access approval Policy-designated resource or data owner Authorized identity/access operator provisions approved access Security or access governance handles policy exceptions
Application incident Application service owner for follow-up Application responder; SRE assists within its agreed remit Incident commander coordinates cross-team response
Platform incident Platform service owner for follow-up Platform responder; CloudOps or SRE assists as agreed Incident commander coordinates affected teams
Reliability objectives Application service owner Application team and SRE develop measurements and engineering work Business stakeholders participate in target decisions
Workload spending Workload budget owner Application team investigates demand; CloudOps supports usage analysis Finance supports forecasting and variance review
Shared platform spending Platform budget owner Platform team investigates usage and allocation Finance and consuming teams review allocation disputes

The important boundary is between providing a capability and using it to operate a workload. AWS’s enablement model describes shared platform services alongside application-team operational ownership. It does not require the platform team to own every application deployed through those services. (docs.aws.amazon.com)

Likewise, SRE involvement should be explicit rather than assumed. Google’s workbook states that product development teams own their services by default in Google’s model, with shared ownership when SRE engages. That is a useful reference model, not a rule every organization must copy. (sre.google)

For spending, distinguish the person authorized to make budget decisions from engineers able to change resource consumption. AWS recommends ongoing collaboration among finance, technology and business stakeholders. (docs.aws.amazon.com)

Close-up of rack-mounted server equipment in a Wikimedia Foundation server rack.
Wikimedia Foundation server equipment photographed in 2012. — Victorgrigas. Own work Source CC BY-SA 3.0

5. Define escalation before an incident crosses team boundaries

An incident needs coordination even when the faulty component’s owner is uncertain. Separate the temporary response roles from long-term service ownership.

Google’s incident-response model identifies three main roles:

  • Incident commander: coordinates the response and delegates responsibilities.
  • Operations lead: directs mitigation and repair work.
  • Communications lead: provides updates and manages stakeholder inquiries.

These roles can expand or contract with the incident; a small response need not assign three separate people immediately. (sre.google)

As a proposed local procedure, define two escalation paths:

Technical escalation: responder → backup → relevant dependency team or provider support.

Ownership escalation: responder or team lead → named authority who can assign temporary responsibility and resolve the boundary.

For handoffs, require an explicit acknowledgment. Include:

  • Current user impact and observed symptoms.
  • Actions already taken and their outcomes.
  • Working hypotheses, clearly distinguished from confirmed facts.
  • Current response roles and the next planned action.

Do not treat an unanswered message as a completed transfer. Set acknowledgment expectations according to your actual support model, rather than inventing a universal timeout.

AWS recommends a discoverable escalation route to someone authorized to assign ownership when responsibility is unclear. (docs.aws.amazon.com)

6. Worked example: a release fails through a shared platform

Consider this hypothetical scenario, not a reported incident:

An application team deploys a change through a shared delivery service. Requests begin failing. The team suspects either the application change or the platform.

Using the larger-organization matrix:

  1. The application responder begins triage. They record customer impact, the release identifier and relevant application evidence. Application failure is not automatically reassigned to platform merely because a shared tool performed the deployment.

  2. The responder attempts the documented mitigation. If rollback is approved for this situation and its prerequisites are satisfied, the application team uses that workflow. If the rollback mechanism fails, it creates a distinct platform investigation.

  3. An incident commander coordinates parallel work. Application engineers investigate application behavior. Platform engineers investigate the failed delivery mechanism. Each workstream reports to the coordinated response.

  4. Access follows the agreed emergency procedure. If additional permissions are necessary, responders use the documented authorization route. An urgent incident is not treated as an undocumented permission grant.

  5. Recovery is checked against user-facing evidence. A successful deployment job alone is not the chosen recovery criterion. In this example, responders verify the affected request path and relevant service indicators.

  6. Follow-up work has separate owners. The application team owns any application correction; platform owns a delivery-service defect; the agreement owner addresses any unclear handoff.

This example combines AWS’s emphasis on safe, reversible change with Google’s coordinated incident roles. It does not assume that rollback is always safe or that a platform defect has been proven. (docs.aws.amazon.com)

7. Write a lightweight agreement and test its weak points

Start with one service and its dependencies. Use this proposed template:

Agreement field What to record
Scope Service, environments and included dependencies
Boundaries What the owner handles and what is excluded
Operational work Release, access, incident, recovery and cost responsibilities
Decision rights Approval authorities and permitted responder actions
Routing Primary contact mechanism, backup and escalation authority
Evidence Runbooks, dashboards, repository and ownership records
Maintenance Agreement owner and review triggers

Keep the document discoverable and assign someone to maintain it. AWS recommends identifiable procedure owners, accessible documentation and mechanisms for review and improvement. (docs.aws.amazon.com)

Then test these common mistakes:

  • “Everyone owns reliability.” Replace the slogan with named decisions, actions and follow-up owners.
  • “The platform deployed it, so platform owns it.” Check the boundary between workload behavior and delivery capability.
  • “SRE owns production.” Specify the services and responsibilities included in the engagement.
  • “Finance owns cost.” Identify both the budget decision-maker and engineers who can investigate consumption.
  • “The runbook author owns the procedure forever.” Verify the current maintainer and a usable backup.
  • “The escalation contact is obvious.” Ask a new responder to find it without private messages.

These checks target gaps highlighted across AWS ownership guidance, Google’s shared-service ownership discussion and AWS’s finance–technology partnership guidance. (docs.aws.amazon.com)

Finally, run a tabletop exercise: an application fails, its owner is unavailable, and a shared dependency may be involved. Check whether participants can find the right route, authorize a mitigation and complete a handoff. AWS recommends simulated failures to test procedures and team response. (docs.aws.amazon.com)

The goal is not perfect terminology. It is an agreement that tells a real responder what they can do, whom to involve and who must resolve the unanswered question.