GitOps: A Comprehensive Guide¶
1. History and Origins¶
1.1 The Kubernetes Problem GitOps Solved¶
By 2015–2016, Kubernetes had become the dominant way to run containerized workloads, but it introduced a new operational problem: cluster state (deployments, services, config maps, ingress rules) was typically applied via imperative kubectl commands run by hand or by ad-hoc scripts from CI jobs. There was no single, reliable, auditable record of what was actually supposed to be running in a cluster, and configuration drift — where the live cluster state silently diverged from what anyone intended — was common and hard to detect.
1.2 Weaveworks Coins the Term¶
In 2017, Alexis Richardson and the team at Weaveworks published a blog post titled "GitOps – Operations by Pull Request," describing how they ran their own Kubernetes-based production systems: all cluster configuration lived in Git, and an automated process continuously reconciled the live cluster to match what was declared in the repository. Rather than pushing changes to production, engineers merged a pull request, and a controller pulled the change and applied it. The name and pattern caught on quickly across the Kubernetes ecosystem.
1.3 Formalization¶
- 2018–2019: Tools purpose-built for the pattern emerge — Flux (from Weaveworks itself) and Argo CD (from Intuit's Argo project) become the two dominant GitOps controllers/operators.
- 2020: The GitOps Working Group forms under the Cloud Native Computing Foundation (CNCF), aiming to standardize terminology and principles across tools.
- 2021: The CNCF's OpenGitOps project publishes the formal GitOps Principles v1.0.0, giving the community a vendor-neutral definition (see below).
- Ongoing: GitOps expands beyond Kubernetes-only use cases into broader infrastructure management, though Kubernetes remains its heartland due to its naturally declarative, reconciliation-friendly API model.
2. Core Principles (OpenGitOps v1.0.0)¶
The CNCF's OpenGitOps working group defines four principles that something must satisfy to be called GitOps:
- Declarative: The desired state of a system is expressed declaratively (YAML manifests, Helm charts, Kustomize overlays, Terraform HCL, etc.) — you describe what the system should look like, not the imperative steps to get there.
- Versioned and Immutable: The desired state is stored in a way that enforces immutability, versioning, and a complete history of changes — in practice, this means Git, giving you commit history, branches, tags, and the ability to revert or roll back trivially.
- Pulled Automatically: Software agents automatically pull the desired state declarations from the source — no human or external CI system pushes changes directly into the runtime environment. This is the inversion at the heart of GitOps: instead of "CI pushes to prod," an in-cluster agent pulls from Git.
- Continuously Reconciled: Software agents continuously observe actual system state and attempt to apply the desired state, correcting drift automatically and alerting when reconciliation isn't possible.
2.1 Push vs. Pull: The Core Distinction from Traditional CI/CD¶
In a traditional CI/CD pipeline, the pipeline itself typically has direct credentials to the production environment and pushes changes into it at the end of a successful run. This means your CI system holds powerful, often broadly-scoped credentials to every environment it deploys to — a significant attack surface.
In GitOps, no external system holds deploy credentials to the cluster. Instead, an agent running inside the target environment (Argo CD, Flux) watches a Git repository and pulls changes when it detects them, then applies them using its own, cluster-local, tightly scoped permissions. The Git repository becomes the single source of truth, and the only way to change production is to change what's declared in Git.
3. Core Components and Architecture¶
A typical GitOps setup involves:
- Source repository (or repositories): Holds the declarative manifests. Many teams split "app config repo" from "app source code repo" so that application code changes and deployment/config changes have independent, auditable histories.
- GitOps operator/controller: Runs inside the cluster (Argo CD, Flux CD) and continuously diffs live state against the Git-declared state.
- Reconciliation loop: The controller pulls the latest manifests, computes the diff against the live cluster, and applies changes to bring the cluster in line — this is the same reconciliation-loop pattern Kubernetes itself is built on (controllers watching for and correcting drift), which is why GitOps fits Kubernetes so naturally.
- Templating/overlay tooling: Helm (templated charts with values files) or Kustomize (base + environment-specific overlays) are commonly used to manage the same underlying manifests across multiple environments (dev/staging/prod) without duplicating YAML.
- Drift detection and alerting: The controller surfaces when the live cluster no longer matches Git — whether due to a manual
kubectlchange, a rogue process, or an external actor — giving visibility that was previously invisible.
4. The GitOps Workflow in Practice¶
- A developer or operator wants to change something — scale a deployment, roll out a new image tag, adjust a config value.
- They open a pull request against the Git repository holding the declarative manifests.
- The PR goes through normal code review, and optionally automated policy checks (e.g., OPA/Gatekeeper policy validation, image signature verification).
- On merge, the GitOps controller detects the change in Git.
- The controller reconciles the cluster to match, applying only the delta.
- If reconciliation fails (e.g., insufficient cluster resources, a bad manifest), the controller reports the failure — often the previous known-good state stays running rather than a broken partial rollout.
- Rolling back is simply reverting the Git commit; the controller reconciles back to the prior state automatically.
This gives you Git's existing tooling — pull requests, required reviewers, commit signing, audit trail, git log, git revert — as your operational control plane for infrastructure, essentially for free.
5. Benefits¶
- Auditability: Every production change has a corresponding Git commit, author, timestamp, and (via PR history) reviewer — invaluable for compliance and incident forensics.
- Reduced credential sprawl: CI systems no longer need broad production credentials; only the in-cluster operator does, and its permissions can be scoped tightly.
- Consistency across environments: The same reconciliation model applies whether you have one cluster or a hundred, and multi-cluster/multi-region fleets can all pull from the same or branched sources of truth.
- Fast, reliable rollback: Reverting a Git commit is a well-understood, low-risk operation compared to ad-hoc rollback scripts.
- Self-healing: Because reconciliation is continuous, manual out-of-band changes (configuration drift) get automatically corrected back to the declared state, rather than silently persisting until someone notices.
6. Common Pitfalls and Criticisms¶
- Secrets management: Git is not a great place to store secrets in plaintext, and GitOps' "everything in Git" philosophy creates real tension here. Solutions include sealed-secrets, external-secrets-operator pulling from a vault/KMS, or SOPS-encrypted files committed to the repo.
- Repository sprawl and structure debates: Mono-repo vs. multi-repo, app-repo vs. config-repo splits, and how to structure overlays for many environments are genuinely difficult design decisions with no universal right answer.
- Not a replacement for CI: GitOps handles the deployment side of the pipeline; you still need CI to build, test, and produce the artifact (container image) whose reference gets updated in the GitOps repo. GitOps and CI/CD are complementary, not substitutes.
- Reconciliation blind spots: If something can change cluster state outside of what the controller manages (e.g., a Horizontal Pod Autoscaler adjusting replica counts), you need to explicitly account for or ignore that field in reconciliation, or you'll get fighting between the autoscaler and the GitOps controller.
- Tooling lock-in and learning curve: Argo CD and Flux have real conceptual and operational differences (Application CRDs vs. Kustomization/HelmRelease CRDs, UI-centric vs. CLI/GitOps-native workflows), and migrating between them isn't trivial once a large fleet is running on one.
7. GitOps in Relation to Other Disciplines¶
- GitOps is best understood as a specific implementation pattern of Continuous Delivery, applying DevOps' broader principles (see the companion DevOps guide) with Git as the explicit, enforced source of truth.
- It pairs naturally with Infrastructure as Code: the declarative manifests GitOps reconciles are IaC by definition; tools like Terraform can also be operated in a GitOps-adjacent style (e.g., Atlantis, tf-controller) even outside Kubernetes.
- It supports DevSecOps goals by making every infrastructure change reviewable and auditable by default, and integrates naturally with policy-as-code admission controllers.
- It's a foundational building block of many Platform Engineering internal developer platforms, since GitOps repos are a natural place to expose a paved-road, self-service interface to application teams.
8. Further Reading¶
- CNCF OpenGitOps project and the formal GitOps Principles (opengitops.dev)
- Argo CD and Flux CD official documentation
- Weaveworks' original "GitOps – Operations by Pull Request" blog post
- GitOps and Kubernetes — Billy Yuen, Alexander Matyushentsev, Todd Ekenstam, Jesse Suen