Every merge to main deploys automatically, infrastructure is reviewable code, and you find out about problems before users tweet them.
Recommended stack
- GitHub Actions - CI/CD. Lives next to the code; matrix builds and environments cover most pipelines.
- Docker - packaging. One artifact that runs identically in CI, staging, and prod.
- Terraform (or OpenTofu) - infra-as-code. Reviewed plans instead of console clicking; state explains what exists and why.
- Grafana + Prometheus/Loki - observability. Metrics, logs, and alerts in one place with sane defaults.
Build steps
- Pipeline order: lint → test → build image → deploy staging → smoke test → promote prod.
- Make deploys boring: immutable image tags, one-command rollback to the previous tag.
- Terraform everything externally visible (DNS, buckets, databases); plan on PR, apply on merge.
- Define 3–5 SLO-based alerts (error rate, latency, saturation); delete alerts nobody acts on.
- Document the break-glass path: how to deploy when CI itself is down.
Watch out for
- Secrets in CI logs or env dumps - use OIDC + a secrets manager.
- Alert fatigue from threshold alerts on every metric.
- Snowflake servers that drift from the Terraform state.
Definition of done
- Merge → production without a human touching a server
- Rollback in under 2 minutes, practiced once
- On-call can diagnose a 500 from dashboards alone