Introduction and guiding principles¶
Good Terraform architecture comes down to one idea: draw boundaries so that each terraform apply touches the smallest set of resources that genuinely change together, and make every environment a different set of inputs to the same code. Everything in this guide (repo layout, state splitting, module design, pipelines) is an application of that idea.
The examples lean on Azure and GitLab CI, but the patterns are provider-agnostic. Where AWS or GCP differ meaningfully, it is called out. Syntax targets Terraform 1.9+ and OpenTofu 1.8+, which share the features used here (moved, import, removed, terraform test, variable validation, optional() object attributes).
The principles that the rest of the guide keeps returning to:
- Same code, different inputs. Dev, staging and prod should differ only in variable values and backend config, never in copied-and-tweaked HCL. Drift between environments' code is how prod surprises you.
- Small blast radius. State files are failure domains. A bad plan, a corrupted state, or a stuck lock should only affect one environment and one layer.
- Separate by rate of change and ownership. Networking changes yearly, apps change daily. Platform teams own the VNet, app teams own their App Service. Put things that change at different speeds, or are owned by different people, in different states.
- Modules encode decisions, not syntax. A good module bakes in your organization's standards (tagging, diagnostics, private endpoints, naming) so consumers can't forget them. A module that just re-exposes every argument of one resource adds a layer and no value.
- Explicit over clever. Readable
for_eachmaps beat nestedcountternaries. A plan a reviewer can understand at 2 a.m. during an incident is worth more than DRYness. - Everything versioned, everything pinned. Terraform core, providers and modules all get pinned versions, and upgrades are deliberate, reviewed changes.
- Humans review, pipelines apply. No one runs
applyagainst shared environments from a laptop. The saved plan that was reviewed is exactly the plan that gets applied.
Core mental model¶
There are three kinds of Terraform code, and most architectural mistakes come from blurring them.
| Kind | What it is | Has state? | Has backend + provider config? | Example |
|---|---|---|---|---|
| Resource module | Wraps one logical resource with standards baked in | No | No | azure-storage-account, aws-s3-bucket |
| Composition (pattern) module | Wires several resource modules into a reusable capability | No | No | web-app = App Service + Key Vault + App Insights + private endpoint |
| Root module (stack) | The thing you plan/apply; instantiates modules for one environment and one layer |
Yes, exactly one state | Yes | live/prod/eastus/network |
The rules that fall out of this¶
- Only root modules configure providers and backends. A child module declares
required_providersbut never contains aproviderblock with credentials or abackendblock. A provider block inside a child module makes it impossible to use withfor_each/countand impossible to remove cleanly. - One root module = one state = one blast radius. If you want two things to fail independently, they need separate roots.
- Environments are instances, not copies.
dev/networkandprod/networkare two root modules that call the same module at (ideally) the same version with different inputs. - Child modules should be environment-agnostic. A module must never contain
if var.env == "prod". Instead, expose the decision as an input (zone_redundant = true,sku = "P1v3") and let the root decide. The moment a module knows environment names, you can't add a new environment without editing it.
Blast radius as the organizing question¶
When deciding where something belongs, ask: if this apply goes wrong, what else breaks? A single state containing the hub VNet, the AKS cluster, and twenty app databases means a typo in an app's tag can produce a plan that touches the network, every plan takes minutes to refresh hundreds of resources, and one team's stuck lock blocks everyone. Splitting by layer and environment keeps plans fast (refresh time scales with resource count), reviews focused, and failures contained.
A reasonable target is a few dozen to a couple hundred resources per state. Above roughly 300-500 resources, plan times and review fatigue usually justify a split. Below ~10, you probably split too eagerly and are now paying a cross-stack wiring tax.
Repository strategy¶
The most durable split is modules repos vs a live repo: modules are versioned libraries, the live repo is the deployed truth for every environment. How you package the modules side is the real choice.
| Strategy | Layout | Strengths | Weaknesses | Fits |
|---|---|---|---|---|
| Monorepo | modules/ and live/ in one repo, modules referenced by relative path |
One MR changes module + callers; simple onboarding; easy grep | No per-module versioning: a module change hits every env on the next apply; promotion is by MR order, not version | Small teams, early stage, a single platform team |
| Live repo + one modules repo | terraform-modules repo with all modules, tagged as a whole (v3.4.0); live repo pins a ref |
Versioned promotion; one release process | A tag bumps every module; changelog noise; one module's breaking change forces a major for all | Mid-size teams with 10-30 modules |
| Live repo + one repo per module | terraform-azurerm-storage-account, etc., each semver-tagged, published to a registry |
Independent versions; clean changelogs; per-module ownership and CI | Many repos to maintain; cross-module changes take several MRs; needs automation (Renovate, a template repo, shared CI) | Larger orgs, multiple consuming teams, a private registry |
Recommendations¶
- Start monorepo, graduate deliberately. Split a module out into its own repo once a second team consumes it or once you need prod to lag dev on that module.
- Keep the live repo boring. It should contain almost no resource blocks, mostly
modulecalls,tfvars, backend config and data lookups. If the live repo is full of rawazurerm_*resources, the modules aren't doing their job. - Use a private registry when you have many module repos. GitLab, Terraform Cloud/HCP, Spacelift and env0 all host one. Registry sources give you
version = "~> 2.1"constraints instead of raw git refs. - Generate new module repos from a template containing
versions.tf,READMEwith terraform-docs markers,examples/,tests/, pre-commit config and a shared CI include. This is what keeps 40 module repos consistent. - Name module repos predictably:
terraform-<provider>-<name>is what public and private registries expect, e.g.terraform-azurerm-key-vault.
A typical monorepo:
infra/
├── modules/
│ ├── resource/ # thin, opinionated wrappers
│ │ ├── key-vault/
│ │ ├── storage-account/
│ │ └── private-endpoint/
│ └── pattern/ # compositions
│ ├── web-app/
│ └── aks-cluster/
├── live/ # one root module per env x region x layer
│ ├── _global/
│ ├── dev/
│ ├── staging/
│ └── prod/
├── policy/ # OPA / Sentinel / Checkov custom rules
├── .tflint.hcl
├── .pre-commit-config.yaml
└── .gitlab-ci.yml
Environment architecture¶
Use one directory per environment × region × layer, each a thin root module calling shared modules. CLI workspaces are the wrong tool for environment separation; Terragrunt or a wrapper is worth it once you have many stacks and the boilerplate starts to hurt.
The four common approaches¶
| Approach | How envs differ | Isolation | Verdict |
|---|---|---|---|
| Directory per env (plain Terraform) | Separate root dirs, each with its own backend and tfvars |
Strong: separate state, can use separate credentials and backends | Default choice. Explicit, greppable, every env visible in the tree. Costs some duplicated backend/provider boilerplate. |
| CLI workspaces | terraform workspace select prod on one root, terraform.workspace in code |
Weak: same backend, same credentials, same code; one forgotten select and you apply dev changes to prod |
Fine for ephemeral copies of one env (per-branch preview stacks). Avoid for dev/stage/prod. |
Single root + -var-file per env |
terraform plan -var-file=prod.tfvars with -backend-config=prod.hcl |
Medium: separate state, but nothing stops the wrong pairing of var-file and backend | Workable with a strict wrapper script or CI-only applies. Easy to fat-finger locally. |
| Terragrunt (or Terramate, Atmos) | terragrunt.hcl per stack inherits shared config, generates backend/provider |
Strong, plus dependency graph and DRY backends | Worth it at roughly 20+ stacks. Adds a tool, a learning curve, and its own failure modes. |
A recommended live layout¶
live/
├── _global/ # once per tenant/account, not per env
│ ├── management-groups/
│ ├── policy-assignments/
│ └── dns-public/
├── shared/ # shared services used by all envs
│ └── canadacentral/
│ ├── hub-network/
│ ├── acr/
│ └── log-analytics/
├── dev/
│ ├── env.hcl / env.auto.tfvars # env-wide values: subscription, tags, sizing tier
│ └── canadacentral/
│ ├── network/ # layer 1: spoke VNet, subnets, NSGs, peering
│ ├── data/ # layer 2: SQL, Storage, Service Bus, Key Vault
│ ├── compute/ # layer 3: AKS / App Service plans
│ └── apps/
│ ├── billing-api/ # layer 4: per-app stacks
│ └── web-frontend/
├── staging/ (same shape)
└── prod/
├── canadacentral/
└── canadaeast/ # DR region
Each leaf directory is a complete root module:
# live/prod/canadacentral/data/main.tf
terraform {
required_version = "~> 1.9"
required_providers {
azurerm = { source = "hashicorp/azurerm", version = "~> 4.10" }
}
backend "azurerm" {
resource_group_name = "rg-tfstate-prod"
storage_account_name = "sttfstateprod001"
container_name = "tfstate"
key = "canadacentral/data.tfstate"
use_azuread_auth = true
}
}
provider "azurerm" {
features {}
subscription_id = var.subscription_id
}
module "sql" {
source = "app.terraform.io/acme/sql-database/azurerm"
version = "2.3.1"
name = "sql-billing-${var.env}-cc"
resource_group_name = data.azurerm_resource_group.data.name
sku_name = var.sql_sku # "GP_S_Gen5_2" in dev, "BC_Gen5_8" in prod
zone_redundant = var.zone_redundant # false in dev, true in prod
subnet_id = data.terraform_remote_state.network.outputs.subnet_ids["data"]
tags = local.tags
}
Keeping environment differences honest¶
- Differences live only in
.tfvars(orenv.hcl), never in conditional code. A reviewer should be able todiff dev/terraform.tfvars prod/terraform.tfvarsand see every intentional difference. - Express differences as capabilities, not names:
zone_redundant,min_replicas,enable_purge_protection,sku. Notis_prod. - Keep a sizing tier map if many values move together:
locals {
tiers = {
small = { app_sku = "B1", sql_sku = "GP_S_Gen5_2", replicas = 1 }
large = { app_sku = "P1v3", sql_sku = "BC_Gen5_8", replicas = 3 }
}
tier = local.tiers[var.tier]
}
- Isolate environments at the account/subscription level, not just the resource-group level. Separate Azure subscriptions (or AWS accounts, GCP projects) per environment give you hard RBAC, quota and billing boundaries, and let each pipeline identity be scoped to exactly one environment.
- Separate state storage per environment. Prod state in a prod-subscription storage account the dev pipeline identity cannot read. State contains secrets in plaintext; treat it accordingly.
Ephemeral environments¶
For per-MR preview stacks, workspaces or a templated key (key = "preview/${MR_IID}.tfstate") are appropriate. Give them a TTL tag, a scheduled destroy job, and a dedicated low-quota subscription so a forgotten preview can't cost real money or touch shared resources.
State architecture¶
Split state into layers that depend in one direction only: foundation → network → data → compute → apps. Lower layers publish outputs; higher layers consume them and never the reverse.
flowchart TB
apps["4 · Applications<br/>Per-app identity, database, queues, app settings, DNS records<br/><i>changes daily</i>"]
compute["3 · Compute platform<br/>AKS clusters, App Service plans, container registries<br/><i>changes weekly</i>"]
data["2 · Shared data and security<br/>Key Vaults, shared SQL, storage, Service Bus and Event Hubs<br/><i>changes weekly</i>"]
network["1 · Network<br/>Hub/spoke VNets, subnets, NSGs, private DNS zones, firewall<br/><i>changes monthly</i>"]
foundation["0 · Global foundation<br/>Management groups, policy, state backends, CI identities<br/><i>changes rarely</i>"]
apps -->|reads| compute
compute -->|reads| data
data -->|reads| network
network -->|reads| foundation
Each box is a separate state per environment; the apps layer changes most and sits furthest from anything hard to rebuild.
Backends¶
| Cloud | Backend | Locking | Notes |
|---|---|---|---|
| Azure | azurerm (Blob Storage) |
Native blob lease | Enable versioning + soft delete; use_azuread_auth = true and disable shared keys |
| AWS | s3 |
use_lockfile = true (S3-native, Terraform 1.10+); DynamoDB table on older versions |
Enable bucket versioning, SSE-KMS, block public access |
| GCP | gcs |
Native | Object versioning on; CMEK if required |
| Any | HCP Terraform / Spacelift / env0 / GitLab-managed state | Built in | GitLab's HTTP backend is convenient if you already live in GitLab CI |
Backend hygiene that applies everywhere:
- Bootstrap the state backend separately, in its own tiny root with local state (or created by a script) and then migrated with
terraform init -migrate-state. Don't let the backend manage itself in a way you can't recover. - Version the state bucket/container. It's your undo button when a state gets mangled.
- Restrict access. State holds every secret Terraform ever saw. Pipeline identity gets read/write for its environment; humans get read-only, break-glass write.
- Name keys predictably, mirroring the directory:
prod/canadacentral/data.tfstate. A key derived from the path means you can always find a stack's state.
Layering¶
A typical split, with how often each changes:
| Layer | Contents | Change rate | Typical owner |
|---|---|---|---|
| 0. Global / foundation | Management groups, policy, subscriptions, state backends, CI identities | Rarely | Platform / cloud governance |
| 1. Network | Hub/spoke VNets, subnets, NSGs, route tables, private DNS zones, firewall | Monthly | Platform / network |
| 2. Shared data and security | Key Vaults, shared SQL servers, storage, Service Bus/Event Hubs namespaces | Weekly | Platform + app teams |
| 3. Compute platform | AKS clusters, App Service plans, container registries | Weekly | Platform |
| 4. Applications | Per-app identities, databases, queues, app settings, DNS records | Daily | App teams |
Rules for layers:
- Dependencies flow downward only. Network never reads from apps. If you find a lower layer needing a higher one's value, the resource is in the wrong layer.
- Things that must be destroyed together live together. An app's database, managed identity and role assignments belong in that app's stack.
- Keep long-lived, hard-to-recreate things away from churn. A SQL server with data or a Key Vault with purge protection should not share a state with resources that get recreated often.
- Kubernetes workloads are usually not Terraform's job. Terraform builds the cluster and its identity plumbing; Helm/Argo CD/Flux deploy into it. Using the Helm or Kubernetes provider in the same state as the cluster that hosts it creates a provider-depends-on-resource chicken-and-egg problem.
Passing data between stacks¶
| Method | How | Pros | Cons |
|---|---|---|---|
| Data sources (preferred) | data "azurerm_subnet" "app" { name = ..., virtual_network_name = ... } |
Loose coupling; reads the real world; no access to another state needed | Requires predictable names or tags to look things up |
terraform_remote_state |
Read another stack's outputs | Simple; exact values | Grants read on the entire other state (secrets included); tight coupling to output names |
| Shared config store | Producer writes to Azure App Configuration / SSM Parameter Store / Consul; consumer reads it | Explicit contract; access-controlled per key | One more system; values can go stale |
Terragrunt dependency blocks |
Terragrunt reads the outputs for you | Ergonomic; builds a run-order graph | Terragrunt-only; mocks needed for plans before the dependency exists |
Prefer data sources with a strong naming convention. Reach for terraform_remote_state only within a trust boundary (same team, same environment), and treat a stack's outputs as a public API: rename them with the same care as a module interface.
# Consumer side: look up by convention, not by reading another state
data "azurerm_subnet" "aks" {
name = "snet-aks"
virtual_network_name = "vnet-${var.env}-cc"
resource_group_name = "rg-network-${var.env}-cc"
}
Module design¶
A good module has a small, typed, documented interface, bakes in your standards, and composes cleanly. Design the interface first, as if writing a public API, because every consumer will couple to it.
Standard file layout¶
terraform-azurerm-key-vault/
├── main.tf # resources
├── variables.tf # inputs, typed + validated + described
├── outputs.tf # outputs, described
├── versions.tf # required_version + required_providers (no provider blocks)
├── locals.tf # naming, tag merging, derived values
├── README.md # generated by terraform-docs between markers
├── CHANGELOG.md
├── examples/
│ ├── basic/ # minimal working root module
│ └── complete/ # every feature on
└── tests/
├── unit.tftest.hcl # command = plan, mocked providers
└── integration.tftest.hcl # command = apply against a sandbox
When to write a module (and when not to)¶
Write one when it encodes a decision: mandatory diagnostics settings, private-endpoint-only networking, a naming convention, required tags, secure TLS defaults, RBAC instead of access policies. Don't write one that wraps a single resource and passes through all 40 of its arguments; that's a maintenance burden with no standard in it. Don't nest modules more than about two levels deep (pattern → resource); deeper trees make plans unreadable and refactors painful.
Inputs¶
- Type everything, using
object()withoptional()defaults for grouped settings instead of dozens of flat variables. - Validate at the edge, so mistakes fail at plan with a clear message rather than as a cryptic API error mid-apply.
- Secure defaults, explicit opt-out.
public_network_access_enableddefaults tofalse; a consumer who wants it public has to say so in code that a reviewer sees. - Few required inputs. A good module works with name, resource group, location and tags. Everything else has a sensible default.
- Mark secrets
sensitive = true, and prefer accepting a Key Vault secret ID or generating secrets inside the module over accepting plaintext. - Accept IDs, not names, for cross-references (
subnet_id, notsubnet_name+vnet_name+rg_name). It removes lookups and works across subscriptions.
variable "network" {
description = "Network exposure. Private by default."
type = object({
public_access_enabled = optional(bool, false)
allowed_ip_ranges = optional(list(string), [])
private_endpoint = optional(object({
subnet_id = string
private_dns_zone_ids = list(string)
}))
})
default = {}
validation {
condition = var.network.public_access_enabled || var.network.private_endpoint != null
error_message = "Either enable public access or supply private_endpoint; otherwise the vault is unreachable."
}
}
variable "name" {
type = string
description = "Key Vault name, 3-24 chars, globally unique."
validation {
condition = can(regex("^[a-zA-Z][a-zA-Z0-9-]{1,22}[a-zA-Z0-9]$", var.name))
error_message = "Name must be 3-24 alphanumerics/hyphens, start with a letter."
}
}
Outputs¶
- Output IDs, names and endpoints consumers need, plus the whole resource object for escape hatches (
output "key_vault" { value = azurerm_key_vault.this }), so new consumer needs don't force a module release. - Describe every output and mark secrets
sensitive. - Outputs are API. Removing or renaming one is a breaking (major) change.
Internals¶
- Name the primary resource
this(azurerm_key_vault.this) for one-per-module resources; it makes references andmovedblocks predictable. - Use
for_eachwith stable string keys, notcount, for collections.countkeys by position, so removing item 0 shifts and recreates every other item.countis fine for a 0-or-1 toggle. - Merge tags in one place:
tags = merge(var.tags, local.module_tags), where module tags record the module name and version. - Generate names from a convention local rather than asking consumers to build them, or use a shared naming module (e.g. Azure/naming).
- No provider blocks, no backend blocks, no hard-coded subscription or region. Accept
location, declarerequired_providerswith a minimum version (>= 4.0), and let the root pin exactly. - Use
configuration_aliaseswhen a module genuinely needs two provider instances (hub and spoke subscriptions for VNet peering), and have the root pass them viaproviders = { azurerm.hub = azurerm.hub }. - Use
lifecyclesparingly and deliberately.prevent_destroyon stateful resources in modules is a good guardrail;ignore_changesshould come with a comment explaining what external system owns that attribute. - Add
precondition/postconditionandcheckblocks for invariants that span resources, e.g. a postcondition that a storage account really hasmin_tls_version = "TLS1_2".
Composition¶
Prefer dependency injection over modules that create their own dependencies. A web-app pattern module should accept log_analytics_workspace_id and subnet_id rather than creating its own workspace and VNet. The root wires shared things once and hands them to every consumer, and the module stays reusable across environments that share or don't share those dependencies.
module "billing_api" {
source = "app.terraform.io/acme/web-app/azurerm"
version = "~> 4.2"
name = "billing-api"
env = var.env
service_plan_id = data.azurerm_service_plan.shared.id
log_analytics_workspace_id = data.azurerm_log_analytics_workspace.central.id
subnet_id = data.azurerm_subnet.apps.id
key_vault_id = module.kv.id
tags = local.tags
}
Documentation¶
Generate the inputs/outputs table with terraform-docs in pre-commit so it never drifts, and hand-write only the parts a generator can't: what decisions the module makes for you, what it deliberately doesn't support, and an upgrade note per major version. A runnable examples/basic is worth more than any prose.
Versioning, pinning and distribution¶
Pin loosely in modules, exactly in roots, and promote module versions through environments the same way you promote application builds.
| Thing | In a child module | In a root module |
|---|---|---|
| Terraform core | required_version = ">= 1.9" |
required_version = "~> 1.9.0" (or exact, via .terraform-version / tfenv / mise) |
| Providers | Minimum only: version = ">= 4.0" |
Pessimistic: version = "~> 4.10", exact via committed .terraform.lock.hcl |
| Modules | n/a | Exact version (version = "2.3.1" or ?ref=v2.3.1) in prod; ~> acceptable in dev |
Semantic versioning for modules¶
- Major: removed/renamed variable or output, changed default that alters infrastructure, anything that forces replacement of existing resources, or requires a
moveddance by the consumer. - Minor: new optional input, new output, new feature behind a flag defaulting to off.
- Patch: bug fixes and docs with no plan diff for existing consumers.
The practical test: if a consumer bumps to this version without changing their code, is their plan empty? If not, it's at least minor; if it destroys anything, it's major. Automate releases with conventional commits and semantic-release (or release-please) so the tag, changelog and registry publish happen from the merge.
Promotion¶
Treat a module version like an artifact moving through environments: bump dev to 2.4.0, let it soak, then open the staging and prod bumps. Renovate (or Dependabot) automates this well: configure it to open one MR per environment directory, auto-merge for dev, and require approval for prod. Pin git sources to tags, never to branches; ?ref=main means prod changes whenever someone merges.
Lock files¶
Commit .terraform.lock.hcl in every root module. Generate hashes for every platform that runs Terraform (developer Macs, Linux CI) so init doesn't rewrite it:
terraform providers lock -platform=linux_amd64 -platform=darwin_arm64 -platform=windows_amd64
Don't commit lock files in child module repos; the consumer's root lock wins anyway.
Provider mirrors and caching¶
In CI, set TF_PLUGIN_CACHE_DIR and cache it between jobs, or run a network mirror (provider_installation in .terraformrc) for air-gapped or rate-limited environments. It makes init fast and removes a dependency on the public registry being up during an incident.
Identity, secrets and providers¶
Pipelines authenticate with short-lived federated credentials, one identity per environment, and secrets are referenced rather than passed through Terraform wherever possible.
Pipeline identity¶
- Use OIDC workload identity federation, not stored client secrets. GitLab CI's
id_tokens→ Azure federated credential (ARM_USE_OIDC=true), AWSAssumeRoleWithWebIdentity, or GCP Workload Identity Federation. No long-lived keys in CI variables. - One identity per environment, scoped to that environment's subscription. The dev pipeline physically cannot touch prod.
- Split plan and apply identities where your platform allows: a read-only identity for
planon MRs (including from forks or untrusted branches), and a write identity only usable from protected branches/environments. Bind federated credentials to the branch or environment claim (ref:main,environment:prod). - Least privilege per layer where practical: the apps-layer identity doesn't need rights to modify the hub VNet. A custom role per layer is more work but limits damage from a compromised pipeline.
Secrets¶
- Assume state is readable by an attacker and minimize what lands in it. Anything Terraform reads or creates (generated passwords, connection strings,
random_password) is stored in state in plaintext. - Prefer identities over secrets. Managed identities / IAM roles plus RBAC eliminate most connection strings entirely.
- Have Terraform create the secret container, not the secret value, when an external process can populate it. Or generate the value inside Terraform and write it straight to Key Vault, accepting it's in state.
- Use ephemeral resources and write-only arguments (Terraform 1.10+/1.11+) where your provider supports them: values that are used during apply but never persisted to state or plan.
- Never put secrets in
.tfvarsin git. Pass viaTF_VAR_*from the CI secret store, or read with adatasource from Key Vault/Secrets Manager. sensitive = trueonly hides values from CLI output. It is redaction, not encryption.
Provider configuration patterns¶
- Configure providers only in roots, from variables, so the same code points at any subscription.
- Use aliases for multi-subscription or multi-region stacks (hub peering, DR replicas, a DNS zone in a connectivity subscription), and pass them into modules explicitly.
- Set
default_tags(AWS) or a centrallocal.tags(Azure) at the root so nothing escapes tagging. On Azure, back it up with an Azure Policy that denies or appends missing required tags. - Register resource providers deliberately. On azurerm 4.x, consider
resource_provider_registrations = "none"for least-privilege identities and register providers in the foundation layer instead.
CI/CD pipelines and promotion¶
Every change goes through the same path: static checks, a saved plan posted to the MR, human review, and an apply of that exact plan file from a protected branch, one environment at a time.
flowchart LR
subgraph mr["On the merge request"]
static[Static checks] --> plan[Plan, read-only] --> policy[Policy gate] --> review[Review + merge]
end
subgraph main["On the default branch, saved plan applied unchanged"]
dev[Apply dev] -->|approval| staging[Apply staging] -->|approval| prod[Apply prod]
end
review -->|merged, dev applies automatically| dev
Staging and prod each wait for a human approval on a protected environment; dev applies on merge.
Pipeline stages¶
- Detect changed stacks. Only plan roots whose files (or whose local modules) changed.
git diff --name-only $CI_MERGE_REQUEST_DIFF_BASE_SHAmapped to root dirs, or let Terragruntrun-all --queue-include-dir/ Terramate--changeddo it. - Static checks (fast, no credentials):
terraform fmt -check -recursive,terraform validate,tflint,checkov/trivy config,terraform-docs --output-check. - Plan each changed stack with the read-only identity:
terraform plan -out=tfplan -lock-timeout=5m, thenterraform show -json tfplanfor policy checks and a human-readable summary posted as an MR comment. GitLab can render plans natively in the MR widget via theterraformreport artifact. - Policy gate on the JSON plan: OPA/Conftest, Sentinel, or Checkov. Block forbidden changes (public IPs in prod, deleting stateful resources, missing tags).
- Review and merge. The reviewer reads the plan, not just the diff. Flag any plan containing
destroyorreplaceon stateful resources for an explicit second approval. - Apply from the protected branch with the write identity, using the saved plan artifact. If the state has moved since the plan (someone else applied), the saved plan is rejected as stale; re-plan rather than forcing.
- Promote by repeating for the next environment: dev applies automatically on merge, staging and prod are manual jobs bound to GitLab protected environments with required approvers.
GitLab CI sketch¶
.tf:
image: hashicorp/terraform:1.9
id_tokens:
AZURE_ID_TOKEN: { aud: api://AzureADTokenExchange }
variables:
ARM_USE_OIDC: "true"
TF_IN_AUTOMATION: "true"
TF_PLUGIN_CACHE_DIR: $CI_PROJECT_DIR/.tf-cache
cache: { key: tf-providers, paths: [.tf-cache] }
before_script:
- cd "$STACK"
- terraform init -input=false
plan:prod:data:
extends: .tf
stage: plan
variables: { STACK: live/prod/canadacentral/data, ARM_CLIENT_ID: $PROD_PLAN_CLIENT_ID }
script:
- terraform plan -input=false -lock-timeout=5m -out=tfplan
- terraform show -json tfplan > plan-full.json # for policy checks
- jq '([.resource_changes[]?.change.actions?] | flatten) | {create: (map(select(. == "create")) | length), update: (map(select(. == "update")) | length), delete: (map(select(. == "delete")) | length)}' plan-full.json > plan.json # MR widget summary
artifacts:
paths: ["$STACK/tfplan", "$STACK/.terraform.lock.hcl"]
reports: { terraform: "$STACK/plan.json" }
rules:
- changes: [live/prod/canadacentral/data/**/*]
apply:prod:data:
extends: .tf
stage: apply
needs: [plan:prod:data]
environment: { name: prod }
variables: { STACK: live/prod/canadacentral/data, ARM_CLIENT_ID: $PROD_APPLY_CLIENT_ID }
script:
- terraform apply -input=false tfplan
rules:
- if: $CI_COMMIT_BRANCH == $CI_DEFAULT_BRANCH
changes: [live/prod/canadacentral/data/**/*]
when: manual
In practice, generate these jobs per stack with a dynamic child pipeline (a script that emits YAML for each changed root) rather than hand-writing hundreds of near-identical jobs.
Operating concerns¶
- Concurrency: state locking prevents corruption, but use
resource_groupin GitLab (or equivalent) so two pipelines don't queue applies against the same stack and fail on lock timeouts. - Drift detection: a scheduled pipeline runs
terraform plan -detailed-exitcodeon every stack nightly; exit code 2 means drift and opens an issue or alerts. Decide per finding whether to codify the change or revert it. - Apply-before-merge vs merge-before-apply: tools like Atlantis apply from the MR, then merge, which guarantees main reflects reality but needs strict locking. Apply-after-merge is simpler. Pick one and be consistent.
- Break-glass: document how a human applies locally in an emergency (who has the role, how to take the lock, how to backfill an MR afterward). Unknown break-glass procedures get improvised badly.
- Plan readability: pipe plans through a summarizer (
tf-summarize,terraform show -json | jq) so reviewers see "3 to add, 1 to change, 0 to destroy" with the resource list first, and the full plan collapsible.
Testing and policy¶
Test modules in layers, cheapest first: static analysis on every commit, mocked terraform test plans on every MR, real applies against a sandbox on release.
| Layer | Tools | What it catches | Runs |
|---|---|---|---|
| Formatting and syntax | terraform fmt, terraform validate |
Style, invalid references, type errors | Pre-commit + every pipeline |
| Linting | tflint with the azurerm/aws/google rulesets |
Invalid SKUs/instance types, deprecated arguments, unused declarations, naming | Pre-commit + every pipeline |
| Security/misconfig scanning | Checkov, Trivy (ex-tfsec), KICS | Public storage, missing encryption, open NSGs, missing diagnostics | Every pipeline, on code and on the JSON plan |
| Unit tests | terraform test with command = plan and mock_provider |
Input validation, conditional logic, naming, tag merging, for_each keys |
Every module MR; no cloud credentials needed |
| Integration tests | terraform test with command = apply, or Terratest (Go) |
The resources actually deploy, connect, and behave | On release tags or nightly, in a sandbox subscription, always destroyed after |
| Policy on plans | OPA/Conftest, Sentinel, Checkov custom policies | Org rules on the change: no deletes of stateful resources in prod, approved regions/SKUs only | Plan stage of live repo pipelines |
| Runtime guardrails | Azure Policy, AWS SCPs/Config, GCP Org Policy | Anything that bypasses Terraform | Always, in the cloud |
terraform test examples¶
# tests/unit.tftest.hcl
mock_provider "azurerm" {}
variables {
name = "kv-test-001"
resource_group_name = "rg-test"
location = "canadacentral"
tags = { owner = "platform" }
}
run "private_by_default" {
command = plan
assert {
condition = azurerm_key_vault.this.public_network_access_enabled == false
error_message = "Key Vault must be private unless explicitly opened."
}
}
run "rejects_bad_name" {
command = plan
variables { name = "1-invalid_name" }
expect_failures = [var.name]
}
Integration tests use the same file format without the mock and with command = apply; Terraform destroys what each test file created at the end. Run them in a dedicated, budget-capped sandbox subscription with a cleanup job that sweeps anything tagged by tests and older than a day, because failed runs will leak resources.
Where policy belongs¶
Use plan-time policy for rules about changes (who may delete what, which environments may have public endpoints) and cloud-native policy for rules about the end state, since it also catches portal clicks and other tools. Overlap is fine; Terraform-level checks give fast feedback in the MR, cloud policy is the real enforcement.
Refactoring and lifecycle¶
Do every refactor declaratively in code (moved, import, removed blocks) so it is reviewed in an MR and applied by the pipeline, instead of running terraform state commands by hand.
| Task | Block | Example |
|---|---|---|
| Rename a resource or wrap it in a module | moved |
moved { from = azurerm_key_vault.main to = module.kv.azurerm_key_vault.this } |
Switch count to for_each |
moved per instance |
moved { from = azurerm_subnet.this[0] to = azurerm_subnet.this["aks"] } |
| Adopt existing (click-ops) resources | import |
import { to = azurerm_resource_group.legacy id = "/subscriptions/.../resourceGroups/rg-legacy" } |
| Bulk-adopt with generated config | import + -generate-config-out |
terraform plan -generate-config-out=generated.tf, then clean up and modularize the output |
| Stop managing something without destroying it | removed |
removed { from = azurerm_storage_account.old lifecycle { destroy = false } } |
| Move a resource to a different state/stack | removed (destroy = false) in source + import in destination |
Two MRs, applied in order: import into the new stack first, then remove from the old |
Guidelines¶
- Ship
movedblocks inside modules when a module release renames internals. Consumers get a clean upgrade with an empty plan, and the change can be a minor version instead of a major. Keep the blocks for at least one major version before deleting them. - Plan must show "0 to destroy" for a pure refactor. If a refactor MR's plan shows replacements, stop; something's mapped wrong.
- Splitting a big state is just the cross-stack move pattern at scale: create the new root,
importevery resource with generatedimportblocks, confirm an empty plan, thenremovedwithdestroy = falsefrom the old root. - Keep
terraform state mv/rmfor emergencies, run with a backup (terraform state pull > backup.tfstate) and recorded in an MR afterward. - Plan provider major upgrades as their own change: bump in dev first, read the upgrade guide, fix deprecations, verify an empty plan, then promote. Never bundle a provider major bump with feature work.
- Deprecate modules visibly: add a
checkblock that warns when a deprecated input is set (checks warn, validations fail), note it in the CHANGELOG and README, and give consumers a version window before removal.
Anti-patterns and checklist¶
Most Terraform pain traces back to a short list of avoidable decisions.
| Anti-pattern | Why it hurts | Instead |
|---|---|---|
| One giant state for everything | Slow plans, huge blast radius, lock contention, terrifying reviews | Split by environment and layer |
| Copy-pasted env directories with diverging HCL | Prod behaves differently from what you tested | Same module calls everywhere; differences only in tfvars |
if env == "prod" inside modules |
Modules can't support new envs without edits | Capability inputs (zone_redundant, sku) |
| CLI workspaces for dev/stage/prod | Same credentials and backend; easy to apply to the wrong env | Directory per environment |
| Provider blocks in child modules | Breaks for_each/count on modules; can't remove the module cleanly |
Providers only in roots, passed via providers = {} |
| Pass-through wrapper modules | Extra layer, no standards, constant churn tracking upstream arguments | Use the resource directly, or a module that encodes decisions |
count for collections |
Removing one item recreates the rest | for_each over a map with stable keys |
Module sources on ?ref=main |
Prod changes whenever anyone merges | Pinned tags, promoted per environment |
| Lock file not committed | Different provider builds in CI vs laptops | Commit .terraform.lock.hcl with multi-platform hashes |
| Applying from laptops | No review, no audit trail, credentials sprawl | Pipeline applies of reviewed saved plans |
terraform_remote_state everywhere |
Tight coupling; consumers read secrets in other states | Data sources by naming convention or a config store |
Secrets in tfvars or outputs without sensitive |
Leaked in git, logs, MR comments | CI secret store, Key Vault data sources, managed identities |
ignore_changes = all to silence drift |
Terraform stops managing the resource without telling anyone | Targeted ignore_changes with a comment, or fix the drift source |
| Deeply nested modules (3+ levels) | Unreadable plans, painful moved blocks |
Flat composition: patterns call resources, roots call patterns |
| Helm/Kubernetes providers in the cluster's own state | Provider can't configure before the cluster exists; destroy ordering breaks | Separate stack, or GitOps (Argo CD/Flux) for in-cluster resources |
Architecture checklist¶
Structure
- [ ] Every environment is a set of root modules calling the same modules; diffs between envs are tfvars only
- [ ] Roots are split by env × region × layer, each with a few dozen to a few hundred resources
- [ ] Layers depend downward only; cross-stack values come from data sources or a deliberate contract
- [ ] Environments are separated at the subscription/account level
State
- [ ] Remote backend with locking, versioning, encryption, and per-env access control
- [ ] State keys mirror directory paths
- [ ] State backend bootstrapped and recoverable independently
Modules
- [ ] Typed, validated, described inputs with secure defaults
- [ ] No provider/backend blocks, no environment names, no hard-coded regions
- [ ]
for_eachwith stable keys for collections - [ ] Generated README, runnable examples,
terraform testunit tests - [ ] Semver releases with changelog;
movedblocks shipped for internal renames
Versions
- [ ] Terraform core, providers and modules pinned in roots; lock files committed
- [ ] Module and provider upgrades promoted dev → staging → prod via automated MRs
Delivery and security
- [ ] OIDC federation, one identity per environment, separate plan/apply identities
- [ ] Saved plans posted to MRs, policy-checked, and applied unchanged from protected branches
- [ ] Manual approval gates on staging/prod; extra approval for destroys of stateful resources
- [ ] Nightly drift detection with alerting
- [ ] Cloud-native policy as the backstop for anything outside Terraform
- [ ] Documented break-glass procedure