Outline only. The full write-up is coming.
The hook#
- A static Hugo site is a folder of files. Mine was spread across 3 repos, 3 cloud environments and a pile of secrets.
- One sentence on where this lands: 2 repos, 1 Pulumi stack, no cloud credentials in CI.
- The thesis, up front: I deleted more than I wrote.
What I built (and why it was fun)#
- The old stack, for context (link back to What powers elyclover.com?):
- 3 repos: the Hugo site, the Pulumi infra program, and a set of reusable composite Actions
- Azure Storage static hosting behind Azure Classic CDN
- one resource group, storage account, CDN profile and CI service principal per environment (
dev,stg,prod) - a prod apex certificate imported as a PFX through Key Vault, because Classic CDN could not issue one for the apex
- SOPS-encrypted secrets and key material kept in Git
- Pulumi pushing service principal credentials into GitHub environments
- What it taught me: IaC, least-privilege CI auth, GitOps with Release Please.
- Honest take: great learning project, wrong amount of machinery for a bio and resume site.
The bill that forced the question#
- Microsoft retired Azure Classic CDN and pointed everyone at Front Door.
- Hosting cost went from about $1/month to about $75/month.
- The last deploys were already failing, because the CDN endpoints had moved underneath the Pulumi state.
- Decision point: pay for Front Door, fix the Classic setup, or rethink the whole thing.
- Pull out the lesson: a cost spike is a good prompt to ask what the system is for.
The plan I threw away#
- v1 moved the stack to GCP behind Cloudflare, partly as a multi-cloud showcase (
MIGRATION_PLAN_v1_gcp.md). - Why it died (my words here): more clouds and more repos to solve a hosting problem that did not need either.
- What replaced it: ask what a static site needs (files, a CDN, TLS, DNS) and let one vendor provide all of it.
- What carried over from v1 (my words here): the habit of writing the plan down first.
What it looks like now (v3)#
- Cloudflare Pages project
elyclover, deployed from GitHub Actions with wrangler Direct Upload. - Cloudflare-managed TLS: no PFX, no Key Vault, no SOPS.
- 2 repos: the site and the Pulumi infra. The reusable Azure Actions repo is retired.
- One Pulumi stack (
prod), written in Go, with tests at 80% coverage or better. - Environments (
dev,stg,production) are Pages branches, not separate clouds. - No cloud credentials in CI: Pulumi mints a scoped Cloudflare token and writes it into GitHub.
- Release Please flow unchanged: PR goes to
dev, release PR goes tostg, release goes to the apex. - Preview environments are
noindex. - Rough before and after table: repos, stacks, secrets in CI, monthly cost, lines of infra code.
The argument: agents made the writing cheap#
- Agents make writing IaC cheap. Pulumi Go, tests, workflows, docs: the typing is no longer the scarce part.
- So the signal is no longer “can you write it”. The signal is judgment about what to build.
- Corollary: the cheapest line of infra code is the one you delete, and an agent will happily build the over-engineered version if you ask.
- Example from this migration: the hard part was deciding what to delete, not writing the replacement.
- Counterpoint to address: does cheap code mean sloppy infra? Tests, review gates and human-only steps for anything destructive.
How the migration was run#
- The plan was a single document,
MIGRATION_PLAN.md: 29 work packages, each with a dependency, an owner and a verify step. - One orchestrator session on Opus read the plan, tracked state in a
STATE.mdfile and launched sub-agents. - Model split:
- Haiku: inventory and verification (run the given commands, compare output)
- Sonnet: code, workflows and docs
- Fable: the plan audit and the code and security review
- Guard rails:
- harness permission rules denied destructive and secret-handling commands for every agent
- agents never saw secret values
- the human did every irreversible step: Azure teardown, token creation, nameserver switch, PR merges
- one agent per repo at a time
- Review before execution: an adversarial review of the plan caught real errors before anything ran.
- Where it went sideways (pick two or three):
- a few spec errors in the plan itself surfaced during execution and had to be fixed by the human
- a deny rule blocked a read-only check it was meant to allow
- a gate was merged out of order and verification had to run retroactively
- What worked: reports written to files, so the orchestrator’s context stayed small.
What I would tell someone else#
- Start with what the site needs, then pick the platform.
- Write the plan down and have a second model attack it before you execute.
- Keep humans on the irreversible steps.
- Treat deleting infrastructure as a feature.
- Link the repos: site and infra.
Open items for me before publishing#
- Replace this outline with prose in my own voice.
- Confirm the cost figures and add the current Cloudflare cost.
- Decide whether to link the v1 GCP plan or only describe it.
- Add one diagram: old stack against new stack.
- Mark the PR ready, merge it, then merge the release PR it creates.