Andrew Mercer
on this page

Introduction and guiding principles

Good Terraform architecture comes down to one idea: draw boundaries so that each terraform apply touches the smallest set of resources that genuinely change together, and make every environment a different set of inputs to the same code. Everything in this guide (repo layout, state splitting, module design, pipelines) is an application of that idea.

The examples lean on Azure and GitLab CI, but the patterns are provider-agnostic. Where AWS or GCP differ meaningfully, it is called out. Syntax targets Terraform 1.9+ and OpenTofu 1.8+, which share the features used here (moved, import, removed, terraform test, variable validation, optional() object attributes).

The principles that the rest of the guide keeps returning to:

  1. Same code, different inputs. Dev, staging and prod should differ only in variable values and backend config, never in copied-and-tweaked HCL. Drift between environments' code is how prod surprises you.
  2. Small blast radius. State files are failure domains. A bad plan, a corrupted state, or a stuck lock should only affect one environment and one layer.
  3. Separate by rate of change and ownership. Networking changes yearly, apps change daily. Platform teams own the VNet, app teams own their App Service. Put things that change at different speeds, or are owned by different people, in different states.
  4. Modules encode decisions, not syntax. A good module bakes in your organization's standards (tagging, diagnostics, private endpoints, naming) so consumers can't forget them. A module that just re-exposes every argument of one resource adds a layer and no value.
  5. Explicit over clever. Readable for_each maps beat nested count ternaries. A plan a reviewer can understand at 2 a.m. during an incident is worth more than DRYness.
  6. Everything versioned, everything pinned. Terraform core, providers and modules all get pinned versions, and upgrades are deliberate, reviewed changes.
  7. Humans review, pipelines apply. No one runs apply against shared environments from a laptop. The saved plan that was reviewed is exactly the plan that gets applied.

Core mental model

There are three kinds of Terraform code, and most architectural mistakes come from blurring them.

Kind What it is Has state? Has backend + provider config? Example
Resource module Wraps one logical resource with standards baked in No No azure-storage-account, aws-s3-bucket
Composition (pattern) module Wires several resource modules into a reusable capability No No web-app = App Service + Key Vault + App Insights + private endpoint
Root module (stack) The thing you plan/apply; instantiates modules for one environment and one layer Yes, exactly one state Yes live/prod/eastus/network

The rules that fall out of this

  • Only root modules configure providers and backends. A child module declares required_providers but never contains a provider block with credentials or a backend block. A provider block inside a child module makes it impossible to use with for_each/count and impossible to remove cleanly.
  • One root module = one state = one blast radius. If you want two things to fail independently, they need separate roots.
  • Environments are instances, not copies. dev/network and prod/network are two root modules that call the same module at (ideally) the same version with different inputs.
  • Child modules should be environment-agnostic. A module must never contain if var.env == "prod". Instead, expose the decision as an input (zone_redundant = true, sku = "P1v3") and let the root decide. The moment a module knows environment names, you can't add a new environment without editing it.

Blast radius as the organizing question

When deciding where something belongs, ask: if this apply goes wrong, what else breaks? A single state containing the hub VNet, the AKS cluster, and twenty app databases means a typo in an app's tag can produce a plan that touches the network, every plan takes minutes to refresh hundreds of resources, and one team's stuck lock blocks everyone. Splitting by layer and environment keeps plans fast (refresh time scales with resource count), reviews focused, and failures contained.

A reasonable target is a few dozen to a couple hundred resources per state. Above roughly 300-500 resources, plan times and review fatigue usually justify a split. Below ~10, you probably split too eagerly and are now paying a cross-stack wiring tax.

Repository strategy

The most durable split is modules repos vs a live repo: modules are versioned libraries, the live repo is the deployed truth for every environment. How you package the modules side is the real choice.

Strategy Layout Strengths Weaknesses Fits
Monorepo modules/ and live/ in one repo, modules referenced by relative path One MR changes module + callers; simple onboarding; easy grep No per-module versioning: a module change hits every env on the next apply; promotion is by MR order, not version Small teams, early stage, a single platform team
Live repo + one modules repo terraform-modules repo with all modules, tagged as a whole (v3.4.0); live repo pins a ref Versioned promotion; one release process A tag bumps every module; changelog noise; one module's breaking change forces a major for all Mid-size teams with 10-30 modules
Live repo + one repo per module terraform-azurerm-storage-account, etc., each semver-tagged, published to a registry Independent versions; clean changelogs; per-module ownership and CI Many repos to maintain; cross-module changes take several MRs; needs automation (Renovate, a template repo, shared CI) Larger orgs, multiple consuming teams, a private registry

Recommendations

  • Start monorepo, graduate deliberately. Split a module out into its own repo once a second team consumes it or once you need prod to lag dev on that module.
  • Keep the live repo boring. It should contain almost no resource blocks, mostly module calls, tfvars, backend config and data lookups. If the live repo is full of raw azurerm_* resources, the modules aren't doing their job.
  • Use a private registry when you have many module repos. GitLab, Terraform Cloud/HCP, Spacelift and env0 all host one. Registry sources give you version = "~> 2.1" constraints instead of raw git refs.
  • Generate new module repos from a template containing versions.tf, README with terraform-docs markers, examples/, tests/, pre-commit config and a shared CI include. This is what keeps 40 module repos consistent.
  • Name module repos predictably: terraform-<provider>-<name> is what public and private registries expect, e.g. terraform-azurerm-key-vault.

A typical monorepo:

infra/
├── modules/
│   ├── resource/          # thin, opinionated wrappers
│   │   ├── key-vault/
│   │   ├── storage-account/
│   │   └── private-endpoint/
│   └── pattern/           # compositions
│       ├── web-app/
│       └── aks-cluster/
├── live/                  # one root module per env x region x layer
│   ├── _global/
│   ├── dev/
│   ├── staging/
│   └── prod/
├── policy/                # OPA / Sentinel / Checkov custom rules
├── .tflint.hcl
├── .pre-commit-config.yaml
└── .gitlab-ci.yml

Environment architecture

Use one directory per environment × region × layer, each a thin root module calling shared modules. CLI workspaces are the wrong tool for environment separation; Terragrunt or a wrapper is worth it once you have many stacks and the boilerplate starts to hurt.

The four common approaches

Approach How envs differ Isolation Verdict
Directory per env (plain Terraform) Separate root dirs, each with its own backend and tfvars Strong: separate state, can use separate credentials and backends Default choice. Explicit, greppable, every env visible in the tree. Costs some duplicated backend/provider boilerplate.
CLI workspaces terraform workspace select prod on one root, terraform.workspace in code Weak: same backend, same credentials, same code; one forgotten select and you apply dev changes to prod Fine for ephemeral copies of one env (per-branch preview stacks). Avoid for dev/stage/prod.
Single root + -var-file per env terraform plan -var-file=prod.tfvars with -backend-config=prod.hcl Medium: separate state, but nothing stops the wrong pairing of var-file and backend Workable with a strict wrapper script or CI-only applies. Easy to fat-finger locally.
Terragrunt (or Terramate, Atmos) terragrunt.hcl per stack inherits shared config, generates backend/provider Strong, plus dependency graph and DRY backends Worth it at roughly 20+ stacks. Adds a tool, a learning curve, and its own failure modes.
live/
├── _global/                    # once per tenant/account, not per env
│   ├── management-groups/
│   ├── policy-assignments/
│   └── dns-public/
├── shared/                     # shared services used by all envs
│   └── canadacentral/
│       ├── hub-network/
│       ├── acr/
│       └── log-analytics/
├── dev/
│   ├── env.hcl / env.auto.tfvars   # env-wide values: subscription, tags, sizing tier
│   └── canadacentral/
│       ├── network/            # layer 1: spoke VNet, subnets, NSGs, peering
│       ├── data/               # layer 2: SQL, Storage, Service Bus, Key Vault
│       ├── compute/            # layer 3: AKS / App Service plans
│       └── apps/
│           ├── billing-api/    # layer 4: per-app stacks
│           └── web-frontend/
├── staging/  (same shape)
└── prod/
    ├── canadacentral/
    └── canadaeast/             # DR region

Each leaf directory is a complete root module:

# live/prod/canadacentral/data/main.tf
terraform {
  required_version = "~> 1.9"
  required_providers {
    azurerm = { source = "hashicorp/azurerm", version = "~> 4.10" }
  }
  backend "azurerm" {
    resource_group_name  = "rg-tfstate-prod"
    storage_account_name = "sttfstateprod001"
    container_name       = "tfstate"
    key                  = "canadacentral/data.tfstate"
    use_azuread_auth     = true
  }
}

provider "azurerm" {
  features {}
  subscription_id = var.subscription_id
}

module "sql" {
  source  = "app.terraform.io/acme/sql-database/azurerm"
  version = "2.3.1"

  name                = "sql-billing-${var.env}-cc"
  resource_group_name = data.azurerm_resource_group.data.name
  sku_name            = var.sql_sku          # "GP_S_Gen5_2" in dev, "BC_Gen5_8" in prod
  zone_redundant      = var.zone_redundant   # false in dev, true in prod
  subnet_id           = data.terraform_remote_state.network.outputs.subnet_ids["data"]
  tags                = local.tags
}

Keeping environment differences honest

  • Differences live only in .tfvars (or env.hcl), never in conditional code. A reviewer should be able to diff dev/terraform.tfvars prod/terraform.tfvars and see every intentional difference.
  • Express differences as capabilities, not names: zone_redundant, min_replicas, enable_purge_protection, sku. Not is_prod.
  • Keep a sizing tier map if many values move together:
locals {
  tiers = {
    small = { app_sku = "B1",   sql_sku = "GP_S_Gen5_2", replicas = 1 }
    large = { app_sku = "P1v3", sql_sku = "BC_Gen5_8",   replicas = 3 }
  }
  tier = local.tiers[var.tier]
}
  • Isolate environments at the account/subscription level, not just the resource-group level. Separate Azure subscriptions (or AWS accounts, GCP projects) per environment give you hard RBAC, quota and billing boundaries, and let each pipeline identity be scoped to exactly one environment.
  • Separate state storage per environment. Prod state in a prod-subscription storage account the dev pipeline identity cannot read. State contains secrets in plaintext; treat it accordingly.

Ephemeral environments

For per-MR preview stacks, workspaces or a templated key (key = "preview/${MR_IID}.tfstate") are appropriate. Give them a TTL tag, a scheduled destroy job, and a dedicated low-quota subscription so a forgotten preview can't cost real money or touch shared resources.

State architecture

Split state into layers that depend in one direction only: foundation → network → data → compute → apps. Lower layers publish outputs; higher layers consume them and never the reverse.

flowchart TB
    apps["4 · Applications<br/>Per-app identity, database, queues, app settings, DNS records<br/><i>changes daily</i>"]
    compute["3 · Compute platform<br/>AKS clusters, App Service plans, container registries<br/><i>changes weekly</i>"]
    data["2 · Shared data and security<br/>Key Vaults, shared SQL, storage, Service Bus and Event Hubs<br/><i>changes weekly</i>"]
    network["1 · Network<br/>Hub/spoke VNets, subnets, NSGs, private DNS zones, firewall<br/><i>changes monthly</i>"]
    foundation["0 · Global foundation<br/>Management groups, policy, state backends, CI identities<br/><i>changes rarely</i>"]
    apps -->|reads| compute
    compute -->|reads| data
    data -->|reads| network
    network -->|reads| foundation

Each box is a separate state per environment; the apps layer changes most and sits furthest from anything hard to rebuild.

Backends

Cloud Backend Locking Notes
Azure azurerm (Blob Storage) Native blob lease Enable versioning + soft delete; use_azuread_auth = true and disable shared keys
AWS s3 use_lockfile = true (S3-native, Terraform 1.10+); DynamoDB table on older versions Enable bucket versioning, SSE-KMS, block public access
GCP gcs Native Object versioning on; CMEK if required
Any HCP Terraform / Spacelift / env0 / GitLab-managed state Built in GitLab's HTTP backend is convenient if you already live in GitLab CI

Backend hygiene that applies everywhere:

  • Bootstrap the state backend separately, in its own tiny root with local state (or created by a script) and then migrated with terraform init -migrate-state. Don't let the backend manage itself in a way you can't recover.
  • Version the state bucket/container. It's your undo button when a state gets mangled.
  • Restrict access. State holds every secret Terraform ever saw. Pipeline identity gets read/write for its environment; humans get read-only, break-glass write.
  • Name keys predictably, mirroring the directory: prod/canadacentral/data.tfstate. A key derived from the path means you can always find a stack's state.

Layering

A typical split, with how often each changes:

Layer Contents Change rate Typical owner
0. Global / foundation Management groups, policy, subscriptions, state backends, CI identities Rarely Platform / cloud governance
1. Network Hub/spoke VNets, subnets, NSGs, route tables, private DNS zones, firewall Monthly Platform / network
2. Shared data and security Key Vaults, shared SQL servers, storage, Service Bus/Event Hubs namespaces Weekly Platform + app teams
3. Compute platform AKS clusters, App Service plans, container registries Weekly Platform
4. Applications Per-app identities, databases, queues, app settings, DNS records Daily App teams

Rules for layers:

  • Dependencies flow downward only. Network never reads from apps. If you find a lower layer needing a higher one's value, the resource is in the wrong layer.
  • Things that must be destroyed together live together. An app's database, managed identity and role assignments belong in that app's stack.
  • Keep long-lived, hard-to-recreate things away from churn. A SQL server with data or a Key Vault with purge protection should not share a state with resources that get recreated often.
  • Kubernetes workloads are usually not Terraform's job. Terraform builds the cluster and its identity plumbing; Helm/Argo CD/Flux deploy into it. Using the Helm or Kubernetes provider in the same state as the cluster that hosts it creates a provider-depends-on-resource chicken-and-egg problem.

Passing data between stacks

Method How Pros Cons
Data sources (preferred) data "azurerm_subnet" "app" { name = ..., virtual_network_name = ... } Loose coupling; reads the real world; no access to another state needed Requires predictable names or tags to look things up
terraform_remote_state Read another stack's outputs Simple; exact values Grants read on the entire other state (secrets included); tight coupling to output names
Shared config store Producer writes to Azure App Configuration / SSM Parameter Store / Consul; consumer reads it Explicit contract; access-controlled per key One more system; values can go stale
Terragrunt dependency blocks Terragrunt reads the outputs for you Ergonomic; builds a run-order graph Terragrunt-only; mocks needed for plans before the dependency exists

Prefer data sources with a strong naming convention. Reach for terraform_remote_state only within a trust boundary (same team, same environment), and treat a stack's outputs as a public API: rename them with the same care as a module interface.

# Consumer side: look up by convention, not by reading another state
data "azurerm_subnet" "aks" {
  name                 = "snet-aks"
  virtual_network_name = "vnet-${var.env}-cc"
  resource_group_name  = "rg-network-${var.env}-cc"
}

Module design

A good module has a small, typed, documented interface, bakes in your standards, and composes cleanly. Design the interface first, as if writing a public API, because every consumer will couple to it.

Standard file layout

terraform-azurerm-key-vault/
├── main.tf            # resources
├── variables.tf       # inputs, typed + validated + described
├── outputs.tf         # outputs, described
├── versions.tf        # required_version + required_providers (no provider blocks)
├── locals.tf          # naming, tag merging, derived values
├── README.md          # generated by terraform-docs between markers
├── CHANGELOG.md
├── examples/
│   ├── basic/         # minimal working root module
│   └── complete/      # every feature on
└── tests/
    ├── unit.tftest.hcl          # command = plan, mocked providers
    └── integration.tftest.hcl   # command = apply against a sandbox

When to write a module (and when not to)

Write one when it encodes a decision: mandatory diagnostics settings, private-endpoint-only networking, a naming convention, required tags, secure TLS defaults, RBAC instead of access policies. Don't write one that wraps a single resource and passes through all 40 of its arguments; that's a maintenance burden with no standard in it. Don't nest modules more than about two levels deep (pattern → resource); deeper trees make plans unreadable and refactors painful.

Inputs

  • Type everything, using object() with optional() defaults for grouped settings instead of dozens of flat variables.
  • Validate at the edge, so mistakes fail at plan with a clear message rather than as a cryptic API error mid-apply.
  • Secure defaults, explicit opt-out. public_network_access_enabled defaults to false; a consumer who wants it public has to say so in code that a reviewer sees.
  • Few required inputs. A good module works with name, resource group, location and tags. Everything else has a sensible default.
  • Mark secrets sensitive = true, and prefer accepting a Key Vault secret ID or generating secrets inside the module over accepting plaintext.
  • Accept IDs, not names, for cross-references (subnet_id, not subnet_name + vnet_name + rg_name). It removes lookups and works across subscriptions.
variable "network" {
  description = "Network exposure. Private by default."
  type = object({
    public_access_enabled = optional(bool, false)
    allowed_ip_ranges     = optional(list(string), [])
    private_endpoint = optional(object({
      subnet_id            = string
      private_dns_zone_ids = list(string)
    }))
  })
  default = {}

  validation {
    condition     = var.network.public_access_enabled || var.network.private_endpoint != null
    error_message = "Either enable public access or supply private_endpoint; otherwise the vault is unreachable."
  }
}

variable "name" {
  type        = string
  description = "Key Vault name, 3-24 chars, globally unique."
  validation {
    condition     = can(regex("^[a-zA-Z][a-zA-Z0-9-]{1,22}[a-zA-Z0-9]$", var.name))
    error_message = "Name must be 3-24 alphanumerics/hyphens, start with a letter."
  }
}

Outputs

  • Output IDs, names and endpoints consumers need, plus the whole resource object for escape hatches (output "key_vault" { value = azurerm_key_vault.this }), so new consumer needs don't force a module release.
  • Describe every output and mark secrets sensitive.
  • Outputs are API. Removing or renaming one is a breaking (major) change.

Internals

  • Name the primary resource this (azurerm_key_vault.this) for one-per-module resources; it makes references and moved blocks predictable.
  • Use for_each with stable string keys, not count, for collections. count keys by position, so removing item 0 shifts and recreates every other item. count is fine for a 0-or-1 toggle.
  • Merge tags in one place: tags = merge(var.tags, local.module_tags), where module tags record the module name and version.
  • Generate names from a convention local rather than asking consumers to build them, or use a shared naming module (e.g. Azure/naming).
  • No provider blocks, no backend blocks, no hard-coded subscription or region. Accept location, declare required_providers with a minimum version (>= 4.0), and let the root pin exactly.
  • Use configuration_aliases when a module genuinely needs two provider instances (hub and spoke subscriptions for VNet peering), and have the root pass them via providers = { azurerm.hub = azurerm.hub }.
  • Use lifecycle sparingly and deliberately. prevent_destroy on stateful resources in modules is a good guardrail; ignore_changes should come with a comment explaining what external system owns that attribute.
  • Add precondition/postcondition and check blocks for invariants that span resources, e.g. a postcondition that a storage account really has min_tls_version = "TLS1_2".

Composition

Prefer dependency injection over modules that create their own dependencies. A web-app pattern module should accept log_analytics_workspace_id and subnet_id rather than creating its own workspace and VNet. The root wires shared things once and hands them to every consumer, and the module stays reusable across environments that share or don't share those dependencies.

module "billing_api" {
  source  = "app.terraform.io/acme/web-app/azurerm"
  version = "~> 4.2"

  name                       = "billing-api"
  env                        = var.env
  service_plan_id            = data.azurerm_service_plan.shared.id
  log_analytics_workspace_id = data.azurerm_log_analytics_workspace.central.id
  subnet_id                  = data.azurerm_subnet.apps.id
  key_vault_id               = module.kv.id
  tags                       = local.tags
}

Documentation

Generate the inputs/outputs table with terraform-docs in pre-commit so it never drifts, and hand-write only the parts a generator can't: what decisions the module makes for you, what it deliberately doesn't support, and an upgrade note per major version. A runnable examples/basic is worth more than any prose.

Versioning, pinning and distribution

Pin loosely in modules, exactly in roots, and promote module versions through environments the same way you promote application builds.

Thing In a child module In a root module
Terraform core required_version = ">= 1.9" required_version = "~> 1.9.0" (or exact, via .terraform-version / tfenv / mise)
Providers Minimum only: version = ">= 4.0" Pessimistic: version = "~> 4.10", exact via committed .terraform.lock.hcl
Modules n/a Exact version (version = "2.3.1" or ?ref=v2.3.1) in prod; ~> acceptable in dev

Semantic versioning for modules

  • Major: removed/renamed variable or output, changed default that alters infrastructure, anything that forces replacement of existing resources, or requires a moved dance by the consumer.
  • Minor: new optional input, new output, new feature behind a flag defaulting to off.
  • Patch: bug fixes and docs with no plan diff for existing consumers.

The practical test: if a consumer bumps to this version without changing their code, is their plan empty? If not, it's at least minor; if it destroys anything, it's major. Automate releases with conventional commits and semantic-release (or release-please) so the tag, changelog and registry publish happen from the merge.

Promotion

Treat a module version like an artifact moving through environments: bump dev to 2.4.0, let it soak, then open the staging and prod bumps. Renovate (or Dependabot) automates this well: configure it to open one MR per environment directory, auto-merge for dev, and require approval for prod. Pin git sources to tags, never to branches; ?ref=main means prod changes whenever someone merges.

Lock files

Commit .terraform.lock.hcl in every root module. Generate hashes for every platform that runs Terraform (developer Macs, Linux CI) so init doesn't rewrite it:

terraform providers lock -platform=linux_amd64 -platform=darwin_arm64 -platform=windows_amd64

Don't commit lock files in child module repos; the consumer's root lock wins anyway.

Provider mirrors and caching

In CI, set TF_PLUGIN_CACHE_DIR and cache it between jobs, or run a network mirror (provider_installation in .terraformrc) for air-gapped or rate-limited environments. It makes init fast and removes a dependency on the public registry being up during an incident.

Identity, secrets and providers

Pipelines authenticate with short-lived federated credentials, one identity per environment, and secrets are referenced rather than passed through Terraform wherever possible.

Pipeline identity

  • Use OIDC workload identity federation, not stored client secrets. GitLab CI's id_tokens → Azure federated credential (ARM_USE_OIDC=true), AWS AssumeRoleWithWebIdentity, or GCP Workload Identity Federation. No long-lived keys in CI variables.
  • One identity per environment, scoped to that environment's subscription. The dev pipeline physically cannot touch prod.
  • Split plan and apply identities where your platform allows: a read-only identity for plan on MRs (including from forks or untrusted branches), and a write identity only usable from protected branches/environments. Bind federated credentials to the branch or environment claim (ref:main, environment:prod).
  • Least privilege per layer where practical: the apps-layer identity doesn't need rights to modify the hub VNet. A custom role per layer is more work but limits damage from a compromised pipeline.

Secrets

  • Assume state is readable by an attacker and minimize what lands in it. Anything Terraform reads or creates (generated passwords, connection strings, random_password) is stored in state in plaintext.
  • Prefer identities over secrets. Managed identities / IAM roles plus RBAC eliminate most connection strings entirely.
  • Have Terraform create the secret container, not the secret value, when an external process can populate it. Or generate the value inside Terraform and write it straight to Key Vault, accepting it's in state.
  • Use ephemeral resources and write-only arguments (Terraform 1.10+/1.11+) where your provider supports them: values that are used during apply but never persisted to state or plan.
  • Never put secrets in .tfvars in git. Pass via TF_VAR_* from the CI secret store, or read with a data source from Key Vault/Secrets Manager.
  • sensitive = true only hides values from CLI output. It is redaction, not encryption.

Provider configuration patterns

  • Configure providers only in roots, from variables, so the same code points at any subscription.
  • Use aliases for multi-subscription or multi-region stacks (hub peering, DR replicas, a DNS zone in a connectivity subscription), and pass them into modules explicitly.
  • Set default_tags (AWS) or a central local.tags (Azure) at the root so nothing escapes tagging. On Azure, back it up with an Azure Policy that denies or appends missing required tags.
  • Register resource providers deliberately. On azurerm 4.x, consider resource_provider_registrations = "none" for least-privilege identities and register providers in the foundation layer instead.

CI/CD pipelines and promotion

Every change goes through the same path: static checks, a saved plan posted to the MR, human review, and an apply of that exact plan file from a protected branch, one environment at a time.

flowchart LR
    subgraph mr["On the merge request"]
        static[Static checks] --> plan[Plan, read-only] --> policy[Policy gate] --> review[Review + merge]
    end
    subgraph main["On the default branch, saved plan applied unchanged"]
        dev[Apply dev] -->|approval| staging[Apply staging] -->|approval| prod[Apply prod]
    end
    review -->|merged, dev applies automatically| dev

Staging and prod each wait for a human approval on a protected environment; dev applies on merge.

Pipeline stages

  1. Detect changed stacks. Only plan roots whose files (or whose local modules) changed. git diff --name-only $CI_MERGE_REQUEST_DIFF_BASE_SHA mapped to root dirs, or let Terragrunt run-all --queue-include-dir / Terramate --changed do it.
  2. Static checks (fast, no credentials): terraform fmt -check -recursive, terraform validate, tflint, checkov/trivy config, terraform-docs --output-check.
  3. Plan each changed stack with the read-only identity: terraform plan -out=tfplan -lock-timeout=5m, then terraform show -json tfplan for policy checks and a human-readable summary posted as an MR comment. GitLab can render plans natively in the MR widget via the terraform report artifact.
  4. Policy gate on the JSON plan: OPA/Conftest, Sentinel, or Checkov. Block forbidden changes (public IPs in prod, deleting stateful resources, missing tags).
  5. Review and merge. The reviewer reads the plan, not just the diff. Flag any plan containing destroy or replace on stateful resources for an explicit second approval.
  6. Apply from the protected branch with the write identity, using the saved plan artifact. If the state has moved since the plan (someone else applied), the saved plan is rejected as stale; re-plan rather than forcing.
  7. Promote by repeating for the next environment: dev applies automatically on merge, staging and prod are manual jobs bound to GitLab protected environments with required approvers.

GitLab CI sketch

.tf:
  image: hashicorp/terraform:1.9
  id_tokens:
    AZURE_ID_TOKEN: { aud: api://AzureADTokenExchange }
  variables:
    ARM_USE_OIDC: "true"
    TF_IN_AUTOMATION: "true"
    TF_PLUGIN_CACHE_DIR: $CI_PROJECT_DIR/.tf-cache
  cache: { key: tf-providers, paths: [.tf-cache] }
  before_script:
    - cd "$STACK"
    - terraform init -input=false

plan:prod:data:
  extends: .tf
  stage: plan
  variables: { STACK: live/prod/canadacentral/data, ARM_CLIENT_ID: $PROD_PLAN_CLIENT_ID }
  script:
    - terraform plan -input=false -lock-timeout=5m -out=tfplan
    - terraform show -json tfplan > plan-full.json   # for policy checks
    - jq '([.resource_changes[]?.change.actions?] | flatten) | {create: (map(select(. == "create")) | length), update: (map(select(. == "update")) | length), delete: (map(select(. == "delete")) | length)}' plan-full.json > plan.json   # MR widget summary
  artifacts:
    paths: ["$STACK/tfplan", "$STACK/.terraform.lock.hcl"]
    reports: { terraform: "$STACK/plan.json" }
  rules:
    - changes: [live/prod/canadacentral/data/**/*]

apply:prod:data:
  extends: .tf
  stage: apply
  needs: [plan:prod:data]
  environment: { name: prod }
  variables: { STACK: live/prod/canadacentral/data, ARM_CLIENT_ID: $PROD_APPLY_CLIENT_ID }
  script:
    - terraform apply -input=false tfplan
  rules:
    - if: $CI_COMMIT_BRANCH == $CI_DEFAULT_BRANCH
      changes: [live/prod/canadacentral/data/**/*]
      when: manual

In practice, generate these jobs per stack with a dynamic child pipeline (a script that emits YAML for each changed root) rather than hand-writing hundreds of near-identical jobs.

Operating concerns

  • Concurrency: state locking prevents corruption, but use resource_group in GitLab (or equivalent) so two pipelines don't queue applies against the same stack and fail on lock timeouts.
  • Drift detection: a scheduled pipeline runs terraform plan -detailed-exitcode on every stack nightly; exit code 2 means drift and opens an issue or alerts. Decide per finding whether to codify the change or revert it.
  • Apply-before-merge vs merge-before-apply: tools like Atlantis apply from the MR, then merge, which guarantees main reflects reality but needs strict locking. Apply-after-merge is simpler. Pick one and be consistent.
  • Break-glass: document how a human applies locally in an emergency (who has the role, how to take the lock, how to backfill an MR afterward). Unknown break-glass procedures get improvised badly.
  • Plan readability: pipe plans through a summarizer (tf-summarize, terraform show -json | jq) so reviewers see "3 to add, 1 to change, 0 to destroy" with the resource list first, and the full plan collapsible.

Testing and policy

Test modules in layers, cheapest first: static analysis on every commit, mocked terraform test plans on every MR, real applies against a sandbox on release.

Layer Tools What it catches Runs
Formatting and syntax terraform fmt, terraform validate Style, invalid references, type errors Pre-commit + every pipeline
Linting tflint with the azurerm/aws/google rulesets Invalid SKUs/instance types, deprecated arguments, unused declarations, naming Pre-commit + every pipeline
Security/misconfig scanning Checkov, Trivy (ex-tfsec), KICS Public storage, missing encryption, open NSGs, missing diagnostics Every pipeline, on code and on the JSON plan
Unit tests terraform test with command = plan and mock_provider Input validation, conditional logic, naming, tag merging, for_each keys Every module MR; no cloud credentials needed
Integration tests terraform test with command = apply, or Terratest (Go) The resources actually deploy, connect, and behave On release tags or nightly, in a sandbox subscription, always destroyed after
Policy on plans OPA/Conftest, Sentinel, Checkov custom policies Org rules on the change: no deletes of stateful resources in prod, approved regions/SKUs only Plan stage of live repo pipelines
Runtime guardrails Azure Policy, AWS SCPs/Config, GCP Org Policy Anything that bypasses Terraform Always, in the cloud

terraform test examples

# tests/unit.tftest.hcl
mock_provider "azurerm" {}

variables {
  name                = "kv-test-001"
  resource_group_name = "rg-test"
  location            = "canadacentral"
  tags                = { owner = "platform" }
}

run "private_by_default" {
  command = plan
  assert {
    condition     = azurerm_key_vault.this.public_network_access_enabled == false
    error_message = "Key Vault must be private unless explicitly opened."
  }
}

run "rejects_bad_name" {
  command = plan
  variables { name = "1-invalid_name" }
  expect_failures = [var.name]
}

Integration tests use the same file format without the mock and with command = apply; Terraform destroys what each test file created at the end. Run them in a dedicated, budget-capped sandbox subscription with a cleanup job that sweeps anything tagged by tests and older than a day, because failed runs will leak resources.

Where policy belongs

Use plan-time policy for rules about changes (who may delete what, which environments may have public endpoints) and cloud-native policy for rules about the end state, since it also catches portal clicks and other tools. Overlap is fine; Terraform-level checks give fast feedback in the MR, cloud policy is the real enforcement.

Refactoring and lifecycle

Do every refactor declaratively in code (moved, import, removed blocks) so it is reviewed in an MR and applied by the pipeline, instead of running terraform state commands by hand.

Task Block Example
Rename a resource or wrap it in a module moved moved { from = azurerm_key_vault.main to = module.kv.azurerm_key_vault.this }
Switch count to for_each moved per instance moved { from = azurerm_subnet.this[0] to = azurerm_subnet.this["aks"] }
Adopt existing (click-ops) resources import import { to = azurerm_resource_group.legacy id = "/subscriptions/.../resourceGroups/rg-legacy" }
Bulk-adopt with generated config import + -generate-config-out terraform plan -generate-config-out=generated.tf, then clean up and modularize the output
Stop managing something without destroying it removed removed { from = azurerm_storage_account.old lifecycle { destroy = false } }
Move a resource to a different state/stack removed (destroy = false) in source + import in destination Two MRs, applied in order: import into the new stack first, then remove from the old

Guidelines

  • Ship moved blocks inside modules when a module release renames internals. Consumers get a clean upgrade with an empty plan, and the change can be a minor version instead of a major. Keep the blocks for at least one major version before deleting them.
  • Plan must show "0 to destroy" for a pure refactor. If a refactor MR's plan shows replacements, stop; something's mapped wrong.
  • Splitting a big state is just the cross-stack move pattern at scale: create the new root, import every resource with generated import blocks, confirm an empty plan, then removed with destroy = false from the old root.
  • Keep terraform state mv/rm for emergencies, run with a backup (terraform state pull > backup.tfstate) and recorded in an MR afterward.
  • Plan provider major upgrades as their own change: bump in dev first, read the upgrade guide, fix deprecations, verify an empty plan, then promote. Never bundle a provider major bump with feature work.
  • Deprecate modules visibly: add a check block that warns when a deprecated input is set (checks warn, validations fail), note it in the CHANGELOG and README, and give consumers a version window before removal.

Anti-patterns and checklist

Most Terraform pain traces back to a short list of avoidable decisions.

Anti-pattern Why it hurts Instead
One giant state for everything Slow plans, huge blast radius, lock contention, terrifying reviews Split by environment and layer
Copy-pasted env directories with diverging HCL Prod behaves differently from what you tested Same module calls everywhere; differences only in tfvars
if env == "prod" inside modules Modules can't support new envs without edits Capability inputs (zone_redundant, sku)
CLI workspaces for dev/stage/prod Same credentials and backend; easy to apply to the wrong env Directory per environment
Provider blocks in child modules Breaks for_each/count on modules; can't remove the module cleanly Providers only in roots, passed via providers = {}
Pass-through wrapper modules Extra layer, no standards, constant churn tracking upstream arguments Use the resource directly, or a module that encodes decisions
count for collections Removing one item recreates the rest for_each over a map with stable keys
Module sources on ?ref=main Prod changes whenever anyone merges Pinned tags, promoted per environment
Lock file not committed Different provider builds in CI vs laptops Commit .terraform.lock.hcl with multi-platform hashes
Applying from laptops No review, no audit trail, credentials sprawl Pipeline applies of reviewed saved plans
terraform_remote_state everywhere Tight coupling; consumers read secrets in other states Data sources by naming convention or a config store
Secrets in tfvars or outputs without sensitive Leaked in git, logs, MR comments CI secret store, Key Vault data sources, managed identities
ignore_changes = all to silence drift Terraform stops managing the resource without telling anyone Targeted ignore_changes with a comment, or fix the drift source
Deeply nested modules (3+ levels) Unreadable plans, painful moved blocks Flat composition: patterns call resources, roots call patterns
Helm/Kubernetes providers in the cluster's own state Provider can't configure before the cluster exists; destroy ordering breaks Separate stack, or GitOps (Argo CD/Flux) for in-cluster resources

Architecture checklist

Structure

  • [ ] Every environment is a set of root modules calling the same modules; diffs between envs are tfvars only
  • [ ] Roots are split by env × region × layer, each with a few dozen to a few hundred resources
  • [ ] Layers depend downward only; cross-stack values come from data sources or a deliberate contract
  • [ ] Environments are separated at the subscription/account level

State

  • [ ] Remote backend with locking, versioning, encryption, and per-env access control
  • [ ] State keys mirror directory paths
  • [ ] State backend bootstrapped and recoverable independently

Modules

  • [ ] Typed, validated, described inputs with secure defaults
  • [ ] No provider/backend blocks, no environment names, no hard-coded regions
  • [ ] for_each with stable keys for collections
  • [ ] Generated README, runnable examples, terraform test unit tests
  • [ ] Semver releases with changelog; moved blocks shipped for internal renames

Versions

  • [ ] Terraform core, providers and modules pinned in roots; lock files committed
  • [ ] Module and provider upgrades promoted dev → staging → prod via automated MRs

Delivery and security

  • [ ] OIDC federation, one identity per environment, separate plan/apply identities
  • [ ] Saved plans posted to MRs, policy-checked, and applied unchanged from protected branches
  • [ ] Manual approval gates on staging/prod; extra approval for destroys of stateful resources
  • [ ] Nightly drift detection with alerting
  • [ ] Cloud-native policy as the backstop for anything outside Terraform
  • [ ] Documented break-glass procedure