Andrew Mercer
on this page

1. History and Origins

1.1 The Problem: DevOps Didn't Scale the Way It Promised

DevOps' "you build it, you run it" model (see the companion DevOps guide) asked every product team to own its full lifecycle — CI/CD, infrastructure provisioning, observability, security scanning, on-call. At small scale, with a handful of teams and strong generalist engineers, this worked. As organizations grew to dozens or hundreds of product teams, a predictable failure mode emerged: every team ended up reinventing its own CI pipeline, its own Terraform modules, its own logging setup, each slightly different, each maintained by developers who would rather be writing product features. Cognitive load exploded. The promise of autonomy turned into the burden of redundant, low-leverage infrastructure work repeated hundreds of times across an org.

This became known, especially in writing by Matthew Skelton and Manuel Pais, as a symptom of excessive cognitive load — teams were being asked to hold too much context (application logic and Kubernetes and cloud networking and security compliance) to be effective at any of it.

1.2 Team Topologies

  • 2019: Matthew Skelton and Manuel Pais publish Team Topologies, introducing a vocabulary that became central to Platform Engineering's justification: four fundamental team types (Stream-Aligned, Platform, Enabling, Complicated-Subsystem) and three interaction modes (Collaboration, X-as-a-Service, Facilitating). The Platform team type is defined explicitly as a team that provides a compelling internal product — a set of self-service APIs, tools, and services — to reduce the cognitive load of stream-aligned (product) teams, so they can focus on their actual business domain.

This gave organizations already drifting toward centralizing DevOps capability a formal, well-argued model for why that centralization made sense — not as a return to the old, siloed Ops team, but as a deliberate, product-minded internal service.

1.3 Emergence as a Named Discipline

  • 2020–2022: As Kubernetes complexity and multi-cloud sprawl made the "every team builds its own platform" model increasingly untenable, "Platform Engineering" begins appearing as a distinct job title and team name across the industry, particularly at companies operating substantial Kubernetes fleets.
  • 2022: Backstage, Spotify's open-source developer portal framework (originally built internally at Spotify and donated to the CNCF in 2020), gains significant adoption, giving teams a concrete, extensible tool for building the "internal developer portal" that Platform Engineering discourse increasingly centered on.
  • 2022–2023: Gartner names Platform Engineering a top strategic technology trend, citing it as the mechanism by which organizations intend to industrialize and scale their DevOps investments; the CNCF's Platform Engineering-adjacent projects and the independent "platformengineering.org" community (founded around this period) formalize much of the vocabulary now in common use (Internal Developer Platform, Golden Paths, etc.).
  • 2023 onward: Platform Engineering becomes a mainstream organizational pattern at mid-size and large engineering organizations, with dedicated conferences (PlatformCon), vendor ecosystems (Humanitec, Port, Cortex, and others building "Internal Developer Platform" products), and increasing convergence with DevSecOps and FinOps concerns as platforms absorb more organizational responsibility.

2. Core Principles

2.1 Treat the Platform as a Product, Not a Project

The single most repeated principle in Platform Engineering literature: the platform team should operate with genuine product management discipline — understanding its internal customers (product/stream-aligned teams) as real users, gathering their feedback, measuring adoption and satisfaction, and prioritizing roadmap work accordingly. A platform built without this discipline tends to reflect what the platform team finds technically interesting rather than what actually reduces friction for its users, and adoption suffers as a result.

2.2 Self-Service Over Tickets

The defining interaction mode (borrowed from Team Topologies) is X-as-a-Service: product teams consume the platform's capabilities through well-documented APIs, CLIs, or portals, without needing to open a ticket and wait for a human on another team to act. This is what actually reduces cognitive load and cycle time — a team provisioning a new database or standing up a new service should take minutes via self-service, not days via a request queue.

2.3 Golden Paths (Paved Roads)

A Golden Path is an opinionated, well-supported, documented route to accomplishing a common task — spinning up a new service, deploying to production, setting up monitoring — that represents the platform team's recommended, "blessed" way of doing things. Golden Paths aren't mandatory walls (teams with genuinely different needs can usually deviate), but they're deliberately made to be the easiest path, so that the default choice for most teams most of the time is also the one that's secure, observable, and consistent with organizational standards by construction.

2.4 Internal Developer Platform (IDP)

The IDP is the actual product platform teams build: the combination of tooling, APIs, and documentation (often surfaced through a developer portal like Backstage) that implements the organization's golden paths. A mature IDP typically abstracts over — rather than replaces — the underlying infrastructure primitives (Kubernetes, cloud APIs, CI/CD systems), giving developers a simpler, higher-level interface while the platform team manages the complexity underneath.

2.5 Reducing Cognitive Load, Not Just Centralizing Ops

It's a common misreading to see Platform Engineering as simply "Ops, rebranded, now gatekeeping everything." The actual goal, per Team Topologies' framing, is narrower and more precise: absorb the undifferentiated complexity (the parts of infrastructure and tooling that are the same regardless of what business problem a given product team is solving) so that product teams' cognitive load is spent on their actual domain — not on relearning Kubernetes networking or Terraform module design from scratch.

3. Core Components of an Internal Developer Platform

  • Developer Portal: A single-pane-of-glass UI (Backstage is the dominant open-source option; commercial alternatives include Port, Cortex, and Humanitec's portal) that catalogs services, ownership, documentation, and provides self-service actions ("scaffold a new service," "request a database").
  • Service Catalog / Software Catalog: A structured inventory of every service, its owning team, its dependencies, its on-call rotation, and its documentation — solving the "who owns this and how does it work" problem that grows acute past a few dozen services.
  • Golden Path Templates / Scaffolding: Templated, opinionated starting points for new services (a "create new microservice" action that generates a repo with CI already wired up, standard observability instrumentation pre-included, and security scanning already configured).
  • CI/CD Abstraction Layer: Rather than every team hand-rolling GitLab CI or GitHub Actions workflows, the platform provides shared, versioned, reusable pipeline templates or a higher-level deployment API.
  • Infrastructure Provisioning Layer: Often built on Terraform, Crossplane, or cloud-native APIs, exposed to product teams through a simplified, self-service interface (a "give me a Postgres database" request) rather than requiring them to write and understand the underlying IaC themselves.
  • Observability and Guardrails Baked In: Golden paths typically ship with logging, metrics, and tracing instrumentation, plus security and policy checks, pre-wired — so that "doing it the easy way" and "doing it the compliant, observable way" are the same path, not competing incentives.

4. Platform Engineering and Team Topologies in Practice

Team Topologies identifies four team types relevant here:

  1. Stream-Aligned Teams: Organized around a business domain or user journey; these are the platform's primary customers.
  2. Platform Teams: Provide the internal product (the IDP) that reduces stream-aligned teams' cognitive load.
  3. Enabling Teams: Temporarily embed with stream-aligned teams to help them adopt new capabilities or unblock a specific challenge, then withdraw — distinct from a platform team's ongoing service relationship.
  4. Complicated-Subsystem Teams: Own a piece of the system requiring deep specialist knowledge (e.g., a video-encoding pipeline, a risk-calculation engine) that would overload a generalist stream-aligned team.

Platform teams interact with stream-aligned teams primarily through the X-as-a-Service mode — self-service consumption with minimal ongoing collaboration overhead — which is precisely what distinguishes a platform team from the old, ticket-queue-driven Ops team using a Collaboration-heavy or purely request-driven interaction mode.

5. Measuring Platform Engineering Success

Because a platform is a product, it should be measured like one:

  • Developer/Team Adoption Rate: What proportion of eligible teams and services actually use the golden path vs. going around it?
  • Time-to-First-Deploy: How long does it take a new service, using the platform, to reach a working production deployment for the first time?
  • Developer Satisfaction / Net Promoter Score: Regularly surveyed sentiment from the platform's internal customers — a frequently underused but highly diagnostic metric.
  • Cognitive Load Proxies: Support ticket volume, onboarding time for new engineers, and the DORA metrics (see the companion DevOps guide) at the product-team level, which platform investment should improve indirectly.
  • Golden Path Deviation Rate: How often teams need to go outside the paved road, and why — a high or rising deviation rate is a signal the golden path no longer fits real usage patterns.

6. Common Misconceptions and Pitfalls

  • "Platform Engineering is just renaming the Ops/Infrastructure team." Without the product mindset, self-service delivery mechanism, and golden-path design discipline, a relabeled Ops team is not a platform team — it just changed its name while keeping the same ticket queue.
  • Building the platform nobody asked for. Platform teams with strong technical opinions but weak product discipline often build sophisticated internal tooling that solves problems their internal customers don't actually have, leading to poor adoption and parallel, unofficial workarounds.
  • Mandating adoption before the platform is actually good. Forcing golden-path adoption before the paved road is genuinely easier and more reliable than the alternative breeds resentment and shadow IT; adoption should be earned through being the better option, not just imposed by policy.
  • Ignoring platform team burnout. Platform teams sit at the intersection of every other team's urgent requests and outages; without the same product-prioritization discipline they're meant to bring to their own roadmap, they become an overloaded, reactive bottleneck rather than a force multiplier.
  • Over-abstracting. Hiding too much of the underlying infrastructure can leave product teams unable to reason about or debug their own systems when something goes wrong outside the golden path's assumptions — abstraction needs to preserve enough transparency for teams to self-serve troubleshooting too, not just self-serve happy-path provisioning.
  • No feedback loop. Platforms built and then left static, without regular user research and iteration, drift out of sync with what product teams actually need as the organization and its technology stack evolve.

7. Platform Engineering in Relation to Other Disciplines

  • Platform Engineering is best understood as the organizational and architectural maturation of DevOps (see the companion DevOps guide) — rather than every team independently practicing "you build it, you run it," a dedicated platform team industrializes the undifferentiated parts of that practice into a shared, self-service product.
  • GitOps is frequently the delivery mechanism underneath a platform's self-service deployment capability — a platform's "deploy my service" action often resolves, under the hood, into a pull request against a GitOps-managed repository (see the companion GitOps guide).
  • SRE practices (SLOs, error budgets, toil reduction) are natural capabilities to build directly into the platform — e.g., a platform that automatically provisions SLO dashboards and alerting for every new service reduces the toil SRE literature warns against (see the companion SRE guide).
  • DevSecOps concerns are ideally addressed by baking security scanning, policy-as-code enforcement, and least-privilege defaults directly into golden paths, so that the secure choice and the easy choice are the same choice rather than competing (see the companion DevSecOps guide).
  • FinOps increasingly gets folded into platform capabilities too — cost visibility and budget guardrails surfaced automatically for every service provisioned through the platform, rather than discovered after the fact in a monthly cloud bill review.

8. Further Reading

  • Team Topologies — Matthew Skelton, Manuel Pais
  • PlatformCon talks and the platformengineering.org community resources
  • Backstage documentation (backstage.io) and the CNCF Backstage project
  • Humanitec's "Platform Engineering" resources and annual State of Platform Engineering surveys
  • Platform Engineering on Kubernetes — Mauricio Salatino