Moving to the cloud as an infrastructure engineer? Start with IaC
This post is for people who have plenty of experience building infrastructure, but are new to the cloud (AWS, Azure, or GCP) and have never used IaC (Infrastructure as Code).
When someone like that starts building infrastructure in the cloud, there are two big new things to learn:
- The cloud platform itself — AWS, Azure, GCP, and so on
- IaC, the practice of managing infrastructure as code
This post covers the second one. It walks through how IaC differs from building by hand in the console, why you should use IaC, and which tool to choose.
What IaC is
IaC is the practice of describing and managing infrastructure — servers, networks, storage — as code instead of by hand. Rather than following manual steps and settings, you use code to provision the infrastructure.
An analogy for on-prem engineers: take the server build procedures you used to keep in paper runbooks and work logs, and turn them into code you can actually run — files you manage and review in Git. If you have used configuration management tools like Ansible, IaC is an extension of the same idea.
The risks of building by hand in the console
Every cloud has a web console where you can create resources from the browser (the AWS Management Console, the Azure Portal, the Google Cloud console). You can click your way to virtual machines and networks, but it comes with risks:
- It is hard to track who changed what and when (the record depends on someone’s memory or screenshots)
- Staging and production drift apart easily, which leads to human error
- Rebuilding from scratch after an incident depends on the memory and skill of whoever built it
- Nobody else can review the work
What IaC solves
With IaC, your infrastructure is written down as text files (code), which means you can:
- Track history and review changes in Git, just like application code
- Roll out the same configuration to staging and production in a reproducible way
- Understand a change just by reading the diff
- Rebuild the environment from code after an incident (no more knowledge locked in one person’s head)
In short, the biggest benefit of IaC is that your infrastructure gets code review too.
Comparing IaC tools
IaC tools fall into two broad groups: tools that work across multiple clouds, and each cloud’s own native tools.
| Terraform | Each cloud’s native tools | |
|---|---|---|
| Examples | Terraform (HashiCorp) | AWS: CloudFormation / AWS CDK Azure: ARM templates / Bicep GCP: Infrastructure Manager (formerly Deployment Manager) |
| Language | HCL (declarative) | YAML / JSON / Bicep; CDK uses general-purpose programming languages like TypeScript or Python |
| Supported clouds | AWS / Azure / GCP and many more | That cloud only |
| State management | Your own state file (stored in S3, Azure Storage, Cloud Storage, etc.) | Generally handled by the cloud |
| What you learn | One way of writing that works on any cloud | A different tool to learn for each cloud |
| Good fit when | You work with several clouds, or may migrate later | You will stay on one cloud and value first-party support |
One more note: GCP’s Infrastructure Manager runs Terraform under the hood, so Terraform knowledge carries straight over to GCP as well. That said, the state of each cloud’s tools changes over time — the move from Deployment Manager to Infrastructure Manager is one example — so check the official documentation for the latest information before you start.
Conclusion: I recommend Terraform. You can manage AWS, Azure, and GCP the same way, the declarative syntax is easy to read, and it has a long track record in real-world use. The native tools have the advantage that the cloud handles state for you, but they only work on that one cloud, so moving to another cloud means learning a new tool from scratch.
Things to watch out for when adopting it
- Learning curve: you need to learn how to write HCL, a declarative language, how resources depend on each other, and the
plan→applyworkflow - Managing the state file: with Terraform, you have to design how you store and operate the state file yourselves. The usual place depends on the cloud — S3 on AWS, Azure Storage (Blob) on Azure, Cloud Storage on GCP — and you should set up locking at the same time to prevent concurrent runs
- Drift: changing things by hand in the console creates a gap (drift) between your code and the real environment. Make “don’t touch it by hand” a firm team rule
- Rollback procedures: Terraform sometimes needs manual intervention when an apply fails, so work out your rollback steps in advance
- Differences between clouds: the way you write Terraform is the same everywhere, but the kinds and names of resources (virtual machines, networks, access management, and so on) differ by cloud. You can’t just copy code from one cloud to another as-is
Next steps
- Start by building something small (a storage bucket, a firewall rule) with Terraform
- Decide where to store the state file and how to lock it, based on the cloud you use
- Set up team rules (no direct changes in the console, changes go through pull requests)
- Once you are comfortable, consider automating
applywith CI/CD (Atlantis, HCP Terraform, etc.)