Andrew Mercer
on this page

cloud-init is the de-facto standard for configuring a Linux instance the first time it boots from a pre-built image — on AWS/Azure/GCP, in OpenStack, or on a local KVM/libvirt VM fed a NoCloud seed. It is not an OS installer (nothing gets partitioned or packages selected from scratch the way Kickstart/Preseed do); it configures an already-installed image: users, SSH keys, hostname, network, packages, and arbitrary commands.

The #cloud-config format

#cloud-config
hostname: server01
fqdn: server01.example.com
manage_etc_hosts: true

users:
  - name: admin
    groups: [sudo]
    shell: /bin/bash
    sudo: 'ALL=(ALL) NOPASSWD:ALL'
    ssh_authorized_keys:
      - ssh-ed25519 AAAA...

package_update: true
package_upgrade: true
packages:
  - chrony
  - vim

write_files:
  - path: /etc/motd
    content: |
      Provisioned by cloud-init
    permissions: '0644'

runcmd:
  - systemctl enable chrony
  - [ sh, -c, "echo done > /var/log/cloud-init-custom.log" ]

power_state:
  mode: reboot
  condition: true

The leading #cloud-config line is required — without it cloud-init tries to interpret the content as a script or another supported format instead of YAML config.

Datasources — where cloud-init gets its config

cloud-init detects a datasource appropriate to the platform it's running on: Ec2, Azure, GCE, OpenStack, etc. on real clouds, or NoCloud/NoCloud-net for local/on-prem use (the same mechanism Ubuntu's Autoinstall rides on):

# NoCloud via an attached ISO ("seed" volume) — needs user-data and meta-data files
genisoimage -output seed.iso -volid cidata -joliet -rock user-data meta-data

# NoCloud-net via HTTP — used heavily with PXE
# kernel param: ds=nocloud-net;s=http://<server>/cloud-init/

meta-data (separate from user-data) minimally needs instance-id, which cloud-init uses to decide whether it has already run for this instance and should not re-run its "first boot" modules.

Check what cloud-init actually did

cloud-init status --long                  # done / running / error, and which datasource was used
cloud-init analyze show                   # timing breakdown of each module/stage
sudo less /var/log/cloud-init.log          # detailed log
sudo less /var/log/cloud-init-output.log   # stdout/stderr of runcmd and package operations
cloud-init schema --config-file user-data.yaml    # validate a config file's schema before using it

Re-running cloud-init for testing (on a cloned/templated image)

sudo cloud-init clean --logs      # reset cloud-init's "already ran" state and logs
sudo cloud-init init
sudo cloud-init modules --mode final

Useful when iterating on a user-data file against a VM template without rebuilding the whole image each time — clean is what makes the instance look "unbooted" to cloud-init again.

Relationship to Autoinstall and Ignition

Autoinstall wraps cloud-init's format with extra installer-specific keys (storage, identity) for driving an actual Ubuntu Server install, then hands the plain user-data: block straight to cloud-init for first boot — so this page's syntax is directly reusable inside an autoinstall config. Ignition solves the same first-boot-configuration problem for Fedora CoreOS/RHCOS but with an entirely separate, non-cloud-init-compatible format.