Infrastructure as Code: Terraform, Pulumi, and State Management
Infrastructure as Code (IaC) turns cloud infrastructure into versioned, reviewable, testable artifacts. Instead of clicking through consoles or running ad-hoc CLI commands, you describe your desired state in code, commit it to Git, and let a tool converge reality to match. Done well, IaC eliminates snowflake environments, shortens onboarding from weeks to hours, and makes disaster recovery a matter of running one pipeline. Done poorly, it produces sprawling state files, brittle modules, and 40-minute plans that nobody trusts. This pillar page walks you through the tooling landscape (Terraform, OpenTofu, Pulumi, CloudFormation, CDK), the mechanics of state management, module design patterns, drift detection, policy as code, and testing strategies that keep IaC reliable at scale.
What Infrastructure as Code Actually Solves
Before IaC, operations teams managed servers by logging in and running commands. If a database needed a replica, someone SSH'd into a box, edited config files, restarted services, and updated a wiki page (usually out of date within a week). Two production environments that were "identical" almost never were. Incident postmortems routinely turned up config drift as the root cause.
IaC codifies infrastructure as declarative source files. You write what you want (an S3 bucket with versioning, a VPC with three private subnets, an RDS instance with 14-day backups) and the tool figures out how to create, update, or destroy resources to match. Git becomes the source of truth. Code review becomes the change management process. Pull request diffs show reviewers exactly which security groups are opening, which IAM policies are changing, and which resources are being replaced.
The second thing IaC solves is repeatability. A staging environment that took two engineers three days to build manually can be reproduced in 20 minutes from code. Disaster recovery drills, ephemeral preview environments, and per-tenant deployments all become tractable. If you can describe it once, you can create it a thousand times.
The third benefit is auditability. Every change is a commit, with an author, timestamp, PR link, and reviewer. Compliance auditors who used to demand screenshots of console configurations now accept Git history. SOC 2 and ISO 27001 evidence collection gets an order of magnitude cheaper.
IaC is also a prerequisite for higher-order practices. GitOps workflows for Kubernetes and cloud infrastructure require declarative state in Git. Policy as code needs machine-readable resource definitions. Cost forecasting tools like Infracost need Terraform plans as input. Without IaC, most of the modern DevOps toolchain is off the table.
The IaC Tool Landscape
The IaC market has four practical categories: HCL-based tools (Terraform, OpenTofu), general-purpose language tools (Pulumi, AWS CDK, CDK for Terraform), cloud-native tools (CloudFormation, Azure Resource Manager, Google Deployment Manager), and configuration management with IaC extensions (Ansible, Chef). Each has trade-offs worth understanding before you commit.
Terraform and OpenTofu
Terraform, created by HashiCorp in 2014, is the de facto standard. It uses HCL (HashiCorp Configuration Language), a declarative DSL designed specifically for infrastructure. Terraform's provider ecosystem covers 3,000+ services including every major cloud, SaaS APIs (GitHub, Datadog, PagerDuty), and on-prem systems (vSphere, Nutanix). Its planning model (compare current state to desired state, propose changes, apply) has become the reference design for the entire category.
In August 2023, HashiCorp relicensed Terraform under the Business Source License, restricting commercial use by competitors. The Linux Foundation forked the last MPL-licensed version as OpenTofu. For the vast majority of teams, OpenTofu is a drop-in replacement; state files, modules, and providers work identically. If you're starting fresh or license-sensitive, OpenTofu is worth evaluating. If you're deep in Terraform Cloud or using recent HCL features, Terraform remains the safe path.
Pulumi
Pulumi lets you write infrastructure in TypeScript, Python, Go, C#, or Java. Instead of learning HCL, you use for loops, functions, classes, and package managers you already know. The mental model is the same as Terraform (declarative, plan and apply, state file) but expressed in general-purpose code.
The upside: complex logic (looping over 47 microservices, computing IAM policies from a data source, generating DNS records from a spreadsheet) is far easier in Python than in HCL. Unit tests run in your language's native test framework. IDE support (autocomplete, refactoring, type checking) is significantly better.
The downside: general-purpose languages let you write general-purpose bugs. It's easy to accidentally create infrastructure inside a loop that shouldn't be there, or to hide critical resources behind abstractions that reviewers can't parse. HCL's constraints are a feature, not a bug, for many teams.
CloudFormation and AWS CDK
If you're all-in on AWS, CloudFormation is the native option. Templates are YAML or JSON, state is managed by AWS (no S3 bucket to configure), and drift detection is built in. The downsides: verbose syntax, slower provider updates than Terraform (new AWS services often take months to appear), and no multi-cloud story.
AWS CDK sits on top of CloudFormation. You write TypeScript, Python, Java, or Go, and CDK synthesizes CloudFormation templates. CDK's higher-level constructs (like aws-ecs-patterns.ApplicationLoadBalancedFargateService) bundle dozens of resources into single lines. It's productive for AWS-only shops but locks you in.
Choosing
| Tool | Language | Multi-cloud | State backend | Best fit |
|---|---|---|---|---|
| Terraform | HCL | Yes | Configurable | General-purpose teams, multi-cloud |
| OpenTofu | HCL | Yes | Configurable | License-conscious teams |
| Pulumi | TS/Py/Go/C#/Java | Yes | Configurable | Teams with strong dev culture |
| CloudFormation | YAML/JSON | AWS only | AWS-managed | AWS-only, compliance-heavy |
| AWS CDK | TS/Py/Java/Go | AWS only | AWS-managed | AWS-only, developer-led |
| Azure Bicep | Bicep DSL | Azure only | Azure-managed | Azure-only shops |
For most new projects at most companies, Terraform or OpenTofu is the safe pick. Pulumi is a strong second choice if your team is developer-heavy and comfortable with the trade-offs. CloudFormation and CDK make sense if you're committed to AWS long-term and value native integrations over portability.
Understanding Terraform's Execution Model
Terraform runs in three phases: init, plan, and apply. Understanding what each does prevents most rookie mistakes.
terraform init downloads provider binaries, initializes the backend (where state lives), and downloads modules. It's cheap to re-run and should be idempotent. Init writes to .terraform/ and creates a lock file (.terraform.lock.hcl) pinning provider versions. Commit the lock file to Git; it prevents subtle upgrades from breaking your team.
terraform plan reads the current state, refreshes it against real cloud APIs, compares to your code, and prints a diff. Symbols matter: + creates, - destroys, ~ updates in place, -/+ destroys and recreates (which for a database means data loss). Read every plan carefully. A common failure mode is skimming the summary line ("11 to add, 3 to change, 1 to destroy") without checking which resource is being destroyed.
terraform apply executes the plan. In CI, always run terraform plan -out=plan.tfplan followed by terraform apply plan.tfplan. This ensures the applied plan is exactly the one reviewed, not a re-planned version that might differ if reality changed in between.
Terraform also has terraform destroy (tear everything down), terraform state (manipulate state directly, dangerous), terraform import (bring existing resources under management), and terraform refresh (sync state with reality). The state and import commands are power tools; learn them before you need them in an incident.
The dependency graph is implicit. When resource B references resource A's attribute (security_group_ids = [aws_security_group.a.id]), Terraform builds a DAG and creates A first. Explicit depends_on is available for cases the DAG can't infer (like an IAM policy that must exist before a service can assume it). Overusing depends_on signals a design problem.
State: The Heart of IaC
State is what separates IaC tools from imperative scripts. The state file is a JSON document mapping resource addresses in your code (like aws_instance.web) to real-world identifiers (like i-0abc123...). Without state, Terraform has no idea which S3 bucket named logs-prod belongs to which of your workspaces or whether it exists at all.
State is also where sensitive data leaks. RDS master passwords, generated TLS private keys, API tokens, all end up in state in plaintext (or base64, which is not encryption). Treat state files like production database backups: encrypt at rest, restrict access, audit reads.
Backend types
Terraform supports many backends. The ones that matter:
- Local: state lives in
terraform.tfstatenext to your code. Fine for personal learning, catastrophic for teams. Two engineers running apply simultaneously will corrupt state. - S3 with DynamoDB locking: the industry standard for AWS. S3 stores the state file (with versioning and encryption enabled), DynamoDB provides locks so only one apply runs at a time. Cross-region replication gives you disaster recovery.
- Azure Storage / Google Cloud Storage: equivalent patterns on other clouds.
- Terraform Cloud / Terraform Enterprise / Spacelift / Env0 / Scalr: managed backends with UI, run history, VCS integration, and policy enforcement.
- Consul / etcd / Postgres: less common but supported.
For production, never use local state, never commit state to Git (secrets leak, no locking), and always enable versioning on your backend so you can recover from accidental deletes or corruption.
State file structure
A minimal state file looks like:
{
"version": 4,
"terraform_version": "1.7.4",
"serial": 142,
"lineage": "8a3f...",
"outputs": {},
"resources": [
{
"mode": "managed",
"type": "aws_s3_bucket",
"name": "logs",
"provider": "provider[\"registry.terraform.io/hashicorp/aws\"]",
"instances": [ { "attributes": { "id": "logs-prod-xyz", "arn": "arn:aws:s3:::logs-prod-xyz" } } ]
}
]
}
The serial increments on every write; lineage identifies the state file across its history. If you ever see two state files with the same lineage but different serials being written to, someone bypassed locking and you have a split-brain problem.
State surgery
Occasionally you need to move resources between state files (splitting a monolith), rename resources without destroying them, or remove resources from state without deleting them from the cloud. These operations use terraform state:
terraform state list: enumerate resourcesterraform state show <address>: inspect a resourceterraform state mv <src> <dst>: rename or move within stateterraform state rm <address>: remove from state, resource stays in cloudterraform import <address> <id>: adopt an existing resource
Always back up state before surgery (terraform state pull > backup.tfstate). If you're using S3 versioning, you can also revert to a prior version. If you're not using versioning, fix that today.
Module Design That Scales
A Terraform module is any directory containing .tf files. When you terraform apply in that directory, you're using the "root module". When you reference another directory with a module block, you're consuming a child module. Good module design is the difference between IaC that scales to hundreds of engineers and IaC that becomes a maintenance nightmare in year two.
The three-layer model
Most mature Terraform organizations converge on a three-layer structure:
- Resource modules wrap a single primitive with sane defaults (an S3 bucket with encryption, versioning, and public access blocking enforced). These are small, opinionated, and reusable.
- Service modules compose resource modules to deliver a functional unit (a web service module wires together an ECS task, ALB target group, DNS record, and IAM role).
- Environment configurations instantiate service modules with environment-specific inputs (prod, staging, dev). These are the "root" modules that run in CI.
This layering keeps blast radius small. Changing a resource module's default (say, enabling S3 object lock) rolls out through service modules to environments deliberately, not all at once.
Inputs, outputs, and defaults
Modules communicate through variables (inputs) and outputs. Design principles:
- Required vs optional: Only make an input required if there is no sensible default.
bucket_namemight be required;versioning_enabledshould default totrue. - Type constraints: Always specify
type = string,type = list(string),type = object({ ... }). Untyped variables cause silent errors. - Validation blocks: Use
validationblocks to reject bad inputs early.condition = can(regex("^[a-z0-9-]+$", var.name))catches typos at plan time. - Descriptions: Every variable and output needs a description. It ends up in generated documentation and PR reviews.
- Outputs are the API: Only expose what consumers need. If you expose too much, you'll break consumers when you refactor internals.
Versioning modules
Publish shared modules to a registry (Terraform Registry, private registry in Terraform Cloud, Git tags, S3, or Artifactory) and pin versions:
module "vpc" {
source = "app.terraform.io/acme/vpc/aws"
version = "3.4.1"
cidr_block = "10.42.0.0/16"
azs = ["us-east-1a", "us-east-1b", "us-east-1c"]
}
Follow semantic versioning. Breaking changes (removing a variable, changing a default that forces replacement) bump the major version. Additive changes bump minor. Bug fixes bump patch. Consumers upgrade on their schedule, not yours.
Anti-patterns to avoid
- The god module: a single module that provisions your entire application. It becomes impossible to change one thing without risking everything.
- The passthrough module: a wrapper that exposes every input of the underlying resource. You've added indirection with no abstraction.
- Deep nesting: modules calling modules calling modules more than two levels deep make plans unreadable and errors hard to trace.
- Copy-paste per environment: three near-identical directories for dev, staging, prod. Use one module and three tfvars files instead.
Workspaces, Environments, and Directory Layout
You have three main options for managing multiple environments: Terraform workspaces, directory-per-environment, and separate backends. Each has trade-offs.
Terraform workspaces are named state files within the same backend. terraform workspace new prod creates a prod workspace; switching workspaces changes which state file is used. Pros: minimal duplication. Cons: it's easy to run apply in the wrong workspace (default is default, which nobody names), and workspaces don't isolate credentials or backends. HashiCorp themselves recommend against workspaces for strong environment separation.
Directory-per-environment creates environments/prod/, environments/staging/, environments/dev/, each with its own terraform.tfvars and backend config. Root modules stay thin, calling into shared service modules. Pros: clear separation, per-environment credentials, hard to apply the wrong thing. Cons: some duplication of backend and provider config.
Separate backends per environment goes further, giving each environment its own state bucket (often in its own AWS account). This is the standard for regulated industries: prod state lives in the prod account with prod-only IAM, and there is no way for a dev-account credential to touch prod state.
A common layout:
infra/
modules/
vpc/
ecs-service/
rds-postgres/
environments/
dev/
main.tf
terraform.tfvars
backend.tf
staging/
main.tf
terraform.tfvars
backend.tf
prod/
main.tf
terraform.tfvars
backend.tf
For AWS Organizations, pair this with one AWS account per environment. Blast radius drops to near zero: a misconfigured IAM policy in dev cannot touch prod resources.
Drift Detection and Reconciliation
Drift is when reality diverges from your code. Someone hotfixed a security group in the console, an autoscaler added tags Terraform doesn't know about, an operator manually resized a database during an incident. Drift is inevitable, so plan for it.
Detecting drift
The simplest detection is terraform plan in CI on a schedule (nightly or hourly). If plan shows changes when nobody merged code, you have drift. Post the plan output to Slack or a dashboard.
Tools that make this easier:
- driftctl (deprecated but instructive): compared cloud state to Terraform state and reported orphaned resources (things in your cloud not managed by any state).
- Terraform Cloud drift detection: built-in, runs on a schedule, sends notifications.
- Spacelift, Env0, Scalr: similar drift detection features in third-party platforms.
- AWS Config: independent of Terraform, tracks resource configuration changes and can trigger notifications on any change.
Handling drift
You have three options when you find drift:
- Absorb: someone made a legitimate change; update your code to match reality. Use
terraform importorterraform statecommands as needed. - Correct: reality is wrong; run apply to revert it. Do this only after understanding why the drift happened; overriding a legitimate emergency fix can cause an outage.
- Investigate: drift you can't classify usually points to a process failure. Who has console write access to prod? Should they?
Long term, the goal is to make drift impossible. Use IAM policies to prevent console writes to Terraform-managed resources. In AWS, an SCP can deny non-read actions except from your CI role. In Kubernetes, admission controllers can reject changes not coming from your GitOps controller. Prevention beats detection.
Policy as Code and Guardrails
Once you have IaC, you can enforce policies on infrastructure changes automatically. This is what "shift left" means in practice: catch violations in a PR, not in a production incident.
Sentinel, OPA, and Checkov
- HashiCorp Sentinel integrates with Terraform Cloud/Enterprise. Policies are written in Sentinel's DSL and run against plan output.
- Open Policy Agent (OPA) / Rego: general-purpose policy engine. Conftest is the CLI that runs Rego against Terraform plans (converted to JSON) or Kubernetes manifests.
- Checkov, tfsec, Terrascan, KICS: opinionated static analyzers with hundreds of built-in rules. They catch things like unencrypted S3 buckets, wide-open security groups, and IAM policies with
*actions.
A minimal Checkov run in CI:
- name: Checkov
uses: bridgecrewio/checkov-action@master
with:
directory: infra/
framework: terraform
soft_fail: false
Start with these scanners on soft-fail (report but don't block), triage the findings, add exceptions for false positives, then flip to hard-fail. Trying to fix everything at once creates change fatigue.
CIS benchmarks and compliance frameworks
For regulated environments, map your policies to standards you already need to meet:
- CIS AWS Foundations Benchmark
- NIST 800-53
- PCI DSS
- HIPAA Security Rule
- SOC 2 Common Criteria
Checkov and its peers ship rule packs for each. Enabling them means your audit evidence is now generated by your PR pipeline. When an auditor asks "how do you enforce encryption at rest?", the answer is "here's the policy, here's every PR it ran against, here's the block on the one that tried to violate it".
Testing IaC: From Lint to Integration
IaC without tests behaves exactly like application code without tests: fine until it isn't. Build a test pyramid.
Static analysis (fast, cheap, run on every commit)
terraform fmt: enforces formatting. Zero excuses to skip.terraform validate: catches syntax errors and missing variable declarations.tflint: pluggable linter with AWS/Azure/GCP rulesets. Catches things like invalid instance types, deprecated arguments, and missing required tags.checkov/tfsec: security policy scanning as described above.
Run all four as pre-commit hooks and as CI checks. They add seconds to your feedback loop and catch a huge percentage of bugs.
Unit tests (medium cost, run on PR)
Terraform 1.6 added a native terraform test framework. You write .tftest.hcl files that run plan or apply against a module and assert on outputs:
run "vpc_creates_three_subnets" {
command = plan
variables {
cidr_block = "10.0.0.0/16"
az_count = 3
}
assert {
condition = length(aws_subnet.private) == 3
error_message = "Expected 3 private subnets"
}
}
For Pulumi, unit tests run in your language's native framework (Jest for TypeScript, pytest for Python) using Pulumi's mocks.
Integration tests (slow, expensive, run nightly or pre-release)
Terratest (Go library from Gruntwork) provisions real infrastructure in a sandbox account, runs assertions against it (HTTP requests, SSH commands, AWS API calls), then tears it down. Tests take 10-30 minutes but catch integration bugs that unit tests can't:
func TestVpcModule(t *testing.T) {
opts := &terraform.Options{ TerraformDir: "../modules/vpc" }
defer terraform.Destroy(t, opts)
terraform.InitAndApply(t, opts)
vpcId := terraform.Output(t, opts, "vpc_id")
assert.NotEmpty(t, vpcId)
}
Kitchen-Terraform and inspec-terraform offer similar capabilities with different ergonomics.
End-to-end / smoke tests (run post-deploy)
After applying to a real environment, run smoke tests: is the load balancer returning 200? Does the database accept a test connection? Are metrics flowing to your observability stack? These belong in your CI/CD pipeline design, not in Terraform itself. They validate that IaC delivered a working system, not just a syntactically correct one.
Secrets, Sensitive Data, and Credentials
Never hardcode secrets in Terraform code. Never. Not even in a variable default. Not even temporarily.
Your options, roughly best to worst:
- Cloud-managed secret stores: AWS Secrets Manager, Azure Key Vault, GCP Secret Manager. Reference secrets by ARN; Terraform fetches them at plan time via data sources.
- HashiCorp Vault: works across clouds, supports dynamic secrets (generate a database credential valid for 15 minutes).
- Environment variables:
TF_VAR_db_password=...from your CI's secret store. Fine for CI, awkward locally. - Encrypted
.tfvarsfiles: with SOPS or git-crypt. Better than plaintext, worse than a secret store. - Plaintext in state: state files always contain secrets in plaintext. Encrypt at rest, restrict IAM, and rotate credentials that leak.
Mark sensitive outputs so Terraform redacts them in logs:
output "db_password" {
value = random_password.db.result
sensitive = true
}
Sensitive still appears in state; it just gets ****ed in plan/apply output. Don't confuse that with encryption.
Cost Management and IaC
Every resource you create has a cost. IaC gives you an opportunity to catch expensive changes before they hit the bill.
Infracost parses Terraform plans and estimates monthly cost. Integrated into a PR, it adds a comment like:
+ aws_db_instance.prod
+ Instance (db.r6g.4xlarge) $1,168.20/month
+ Storage (gp3, 500GB) $57.50/month
Total change: +$1,225.70/month
Reviewers see cost impact next to the diff. A PR that accidentally scales up a fleet by 10x becomes obvious. Set thresholds: PRs adding more than $500/month require a second approver or a comment justifying the change.
Beyond estimates, tag everything. Every resource should carry Environment, Team, CostCenter, Application tags. Use default_tags in the AWS provider so you don't repeat yourself:
provider "aws" {
default_tags {
tags = {
Environment = var.environment
ManagedBy = "terraform"
Repo = "acme/infra"
}
}
}
Cost reports segmented by tag turn "we spent $2M on compute" into "the checkout team spent $340K on staging RDS". That specificity drives action.
Multi-Cloud and Provider Patterns
If you truly run multi-cloud (not "we might someday" but "we have production workloads on two clouds"), IaC helps but doesn't magically abstract clouds away. An aws_instance is not a google_compute_instance even if both are VMs. Terraform gives you consistent tooling and workflow; the resources themselves stay cloud-specific.
Patterns that work:
- Provider aliases: run multiple accounts or regions in one config.
provider "aws" { alias = "us-east-1" }andprovider "aws" { alias = "eu-west-1" }let you deploy to both from one apply. - Shared abstractions at the module boundary: a
data_lakemodule might acceptcloud = "aws"orcloud = "gcp"and internally provision the right primitives. This only pays off for a small number of high-value abstractions. - Per-cloud state: don't put AWS and GCP resources in the same state file. Failures in one cloud shouldn't block the other.
Most "multi-cloud" IaC is really "multi-account within one cloud plus a small footprint elsewhere". Design for that reality.
IaC in CI/CD Pipelines
IaC pipelines have a predictable shape:
- On PR open:
terraform fmt -check,terraform validate,tflint,checkov,terraform plan, post plan as PR comment. - On PR merge to main:
terraform plan -out=plan.tfplan, request approval for prod,terraform apply plan.tfplan. - On schedule:
terraform planfor drift detection, alert if changes detected.
Guardrails to add:
- Plan approval: prod applies require human review of the plan, even if the code was reviewed. Plans can change if state changed.
- Concurrency locks: only one apply per environment at a time. Your backend lock handles this, but your CI should also queue jobs so failures are clear.
- Least-privilege CI credentials: your CI role should not have
*:*. Scope it to what it actually manages. In AWS, use IAM Access Analyzer to generate least-privilege policies from CloudTrail history. - Backend state access: only the CI role writes state. Humans get read access for debugging. If humans can write state, humans can corrupt it.
Tools like Atlantis, Terraform Cloud, Spacelift, Env0, and Scalr provide this workflow out of the box. Building it yourself in GitHub Actions or GitLab CI is fine for small teams; the platforms pay off once you have 20+ engineers touching infra.
For deeper context on pipeline design overall, see the CI/CD pipeline guide in this DevOps hub.
GitOps and IaC Together
IaC and GitOps are complementary. IaC covers cloud infrastructure (VPCs, databases, IAM). GitOps typically covers Kubernetes workloads (Deployments, Services, ConfigMaps) via controllers like Argo CD or Flux that continuously reconcile cluster state to Git.
The pattern that scales: Terraform provisions the EKS/GKE/AKS cluster and its bootstrap addons (Argo CD, cert-manager, external-secrets-operator). Argo CD takes over from there, deploying application workloads from another Git repo. Terraform runs weekly or on infra changes; Argo CD reconciles every three minutes.
This split respects each tool's strengths. Terraform is good at slow-changing, dependency-heavy, cloud-primitive provisioning. GitOps controllers are good at fast, continuous, Kubernetes-native reconciliation. The GitOps patterns page covers the Kubernetes side in depth.
Migrating From Clicked-Together Infrastructure
Most teams don't start with IaC; they inherit a cloud account someone built by hand and need to bring it under management. The migration playbook:
- Freeze console changes: announce that from a specific date, all infra changes go through code. Enforce with IAM after a grace period.
- Inventory: enumerate resources.
aws-nuke --dry-run,terraformer,former2, or cloud-native inventory tools give you a list. - Prioritize by blast radius: import your most critical resources first (VPCs, IAM, databases). Import ephemeral things (EC2 instances that autoscale) last or not at all.
- Import incrementally: write the Terraform code for a resource,
terraform import <address> <id>, run plan, expect diffs, iterate until plan is clean. - Reorganize: once imported, refactor into modules. Don't try to import into a beautiful module structure on day one; land it, then improve.
- Prevent regression: SCPs or IAM policies that deny console writes for Terraform-managed resource types.
Expect the first pass to be ugly. That's fine. Ugly IaC beats no IaC. You can refactor once state is under control.
Skills, Career Path, and Learning IaC
IaC is a skill that pays off across every DevOps and platform role. Job titles that require it: DevOps Engineer, Platform Engineer, Site Reliability Engineer, Cloud Engineer, Infrastructure Engineer, Security Engineer. Salary bands are consistently 10-20% above general backend roles at the same level, because supply is tight.
To build practical skill:
- Learn one cloud deeply first: AWS is the safest bet for job market. Get to the point where you can provision a VPC, ECS/EKS service, RDS, and ALB from Terraform without looking things up constantly.
- Build a personal project: a real thing you deploy from code. A blog, a Discord bot, a home lab in AWS. Anything that survives multiple
terraform applycycles and forces you to handle real problems (state, secrets, upgrades). - Read other people's modules: the Terraform AWS Modules organization (
terraform-aws-modules/*) publishes hundreds of high-quality modules. Read them. Notice conventions. - Practice failure: intentionally break things. Delete state. Corrupt a module. Watch what happens. Recover.
- Get certified if it helps you: HashiCorp's Terraform Associate certification is entry-level but signals baseline literacy. AWS/Azure/GCP cloud certs are more impactful.
The Refonte Learning DevOps engineer program covers Terraform, GitOps, CI/CD, Kubernetes, and observability with hands-on labs and an internship component. If you want a structured path from beginner to job-ready, that program is designed for exactly this trajectory, including the IaC depth this page describes.
For a broader view of adjacent topics (containers, observability, incident response), see the DevOps learning hub index.
Common Pitfalls and How to Avoid Them
A few failure modes that catch nearly every team at some point:
- Giant state files: one state file managing 5,000 resources takes 15+ minutes to plan. Split by environment first, then by domain (network, data, compute, apps).
countvsfor_each:countuses list indexes, so removing item 2 from a list of 5 will recreate items 3, 4, 5. Usefor_eachwith a map for stable addressing.- Provider version drift: not pinning provider versions means one engineer's apply differs from another's. Commit
.terraform.lock.hcl. - Ignoring plan output: a
-/+on your production database is not a routine change. Read every plan. - Bypassing locks:
terraform apply -lock=falseexists but is almost never the right answer. If a lock is stuck, understand why before forcing. - Storing secrets in code: even in a private repo, secrets in Git are secrets in every developer's laptop, every CI cache, and every backup forever. Rotate immediately if it happens.
- No disaster recovery for state: if your state bucket is deleted, can you recover? Test it. Cross-region replication, versioning, and periodic backups to a separate account are the belt-and-suspenders setup.
- Manual changes "just this once": they always come back to bite you. If the process is too slow for emergencies, fix the process, don't bypass IaC.
Frequently Asked Questions
Do I need to learn Terraform if my company uses Pulumi (or vice versa)? Learn the concepts (state, plan/apply, modules, drift) once and both tools map cleanly. Syntax is the easy part. If you know Terraform deeply, picking up Pulumi is a weekend. The market has more Terraform jobs, so start there if you're choosing.
Should I use Terraform workspaces for environments? Generally no. Use directory-per-environment with separate backends (ideally separate cloud accounts). Workspaces are useful for short-lived branches or ephemeral test environments, not for isolating dev/staging/prod.
How do I handle Terraform version upgrades?
Upgrade minor versions freely; major versions (0.x to 1.x, or between 1.x majors) may require code changes. Test in a non-prod environment first. Pin the version in required_version and in your CI image so upgrades are deliberate.
Is CDK better than Terraform because it uses TypeScript? Different, not better. CDK is productive for AWS-focused teams that value language-native tooling. Terraform is broader (multi-cloud, larger provider ecosystem, more third-party tooling like Infracost and Checkov). Pick based on your cloud strategy and team composition.
What's the difference between terraform import and adopting resources via config?
terraform import puts an existing cloud resource into state so Terraform knows it exists. You still have to write the corresponding .tf code manually. Terraform 1.5 added import blocks that let you declare imports in code, which is easier to review and version.
How do I test Terraform without spending a fortune in cloud bills?
Static analysis (fmt, validate, tflint, checkov) is free and catches most bugs. Native terraform test with command = plan doesn't create real resources. For integration tests with Terratest, use a dedicated sandbox account with billing alerts, tear down resources aggressively, and run integration tests on merge or nightly, not per commit.
Can I use IaC without Kubernetes? Absolutely. IaC applies to any cloud infrastructure: VMs, serverless functions, databases, DNS, load balancers, IAM, networking. Kubernetes is one workload target among many, and many organizations run substantial infra with no Kubernetes at all.
What should I read next? Depending on your gap: the GitOps deep dive if you're moving to Kubernetes-native workflows, the CI/CD design guide if you need to wire IaC into pipelines, or the DevOps engineer program if you want a structured mentored path from fundamentals through production-grade IaC.
