↓ Skip to main content
  1. Articles/

Why I killed my over-engineered portfolio infrastructure

·5 mins

Outline only. The full write-up is coming.

The hook
#

  • A static Hugo site is a folder of files. Mine was spread across 3 repos, 3 cloud environments and a pile of secrets.
  • One sentence on where this lands: 2 repos, 1 Pulumi stack, no cloud credentials in CI.
  • The thesis, up front: I deleted more than I wrote.

What I built (and why it was fun)
#

  • The old stack, for context (link back to What powers elyclover.com?):
    • 3 repos: the Hugo site, the Pulumi infra program, and a set of reusable composite Actions
    • Azure Storage static hosting behind Azure Classic CDN
    • one resource group, storage account, CDN profile and CI service principal per environment (dev, stg, prod)
    • a prod apex certificate imported as a PFX through Key Vault, because Classic CDN could not issue one for the apex
    • SOPS-encrypted secrets and key material kept in Git
    • Pulumi pushing service principal credentials into GitHub environments
  • What it taught me: IaC, least-privilege CI auth, GitOps with Release Please.
  • Honest take: great learning project, wrong amount of machinery for a bio and resume site.

The bill that forced the question
#

  • Microsoft retired Azure Classic CDN and pointed everyone at Front Door.
  • Hosting cost went from about $1/month to about $75/month.
  • The last deploys were already failing, because the CDN endpoints had moved underneath the Pulumi state.
  • Decision point: pay for Front Door, fix the Classic setup, or rethink the whole thing.
  • Pull out the lesson: a cost spike is a good prompt to ask what the system is for.

The plan I threw away
#

  • v1 moved the stack to GCP behind Cloudflare, partly as a multi-cloud showcase (MIGRATION_PLAN_v1_gcp.md).
  • Why it died (my words here): more clouds and more repos to solve a hosting problem that did not need either.
  • What replaced it: ask what a static site needs (files, a CDN, TLS, DNS) and let one vendor provide all of it.
  • What carried over from v1 (my words here): the habit of writing the plan down first.

What it looks like now (v3)
#

  • Cloudflare Pages project elyclover, deployed from GitHub Actions with wrangler Direct Upload.
  • Cloudflare-managed TLS: no PFX, no Key Vault, no SOPS.
  • 2 repos: the site and the Pulumi infra. The reusable Azure Actions repo is retired.
  • One Pulumi stack (prod), written in Go, with tests at 80% coverage or better.
  • Environments (dev, stg, production) are Pages branches, not separate clouds.
  • No cloud credentials in CI: Pulumi mints a scoped Cloudflare token and writes it into GitHub.
  • Release Please flow unchanged: PR goes to dev, release PR goes to stg, release goes to the apex.
  • Preview environments are noindex.
  • Rough before and after table: repos, stacks, secrets in CI, monthly cost, lines of infra code.

The argument: agents made the writing cheap
#

  • Agents make writing IaC cheap. Pulumi Go, tests, workflows, docs: the typing is no longer the scarce part.
  • So the signal is no longer “can you write it”. The signal is judgment about what to build.
  • Corollary: the cheapest line of infra code is the one you delete, and an agent will happily build the over-engineered version if you ask.
  • Example from this migration: the hard part was deciding what to delete, not writing the replacement.
  • Counterpoint to address: does cheap code mean sloppy infra? Tests, review gates and human-only steps for anything destructive.

How the migration was run
#

  • The plan was a single document, MIGRATION_PLAN.md: 29 work packages, each with a dependency, an owner and a verify step.
  • One orchestrator session on Opus read the plan, tracked state in a STATE.md file and launched sub-agents.
  • Model split:
    • Haiku: inventory and verification (run the given commands, compare output)
    • Sonnet: code, workflows and docs
    • Fable: the plan audit and the code and security review
  • Guard rails:
    • harness permission rules denied destructive and secret-handling commands for every agent
    • agents never saw secret values
    • the human did every irreversible step: Azure teardown, token creation, nameserver switch, PR merges
    • one agent per repo at a time
  • Review before execution: an adversarial review of the plan caught real errors before anything ran.
  • Where it went sideways (pick two or three):
    • a few spec errors in the plan itself surfaced during execution and had to be fixed by the human
    • a deny rule blocked a read-only check it was meant to allow
    • a gate was merged out of order and verification had to run retroactively
  • What worked: reports written to files, so the orchestrator’s context stayed small.

What I would tell someone else
#

  • Start with what the site needs, then pick the platform.
  • Write the plan down and have a second model attack it before you execute.
  • Keep humans on the irreversible steps.
  • Treat deleting infrastructure as a feature.
  • Link the repos: site and infra.

Open items for me before publishing
#

  • Replace this outline with prose in my own voice.
  • Confirm the cost figures and add the current Cloudflare cost.
  • Decide whether to link the v1 GCP plan or only describe it.
  • Add one diagram: old stack against new stack.
  • Mark the PR ready, merge it, then merge the release PR it creates.