The Petone Exit: Decommissioning

The colo exit deadline arrives. Reach matrices, decommissioned management groups, and leaving as the real LZ test. Written for NZ platform teams.

Part 10
SR
Steve Rackham
14 min read Guides

Part 10 of “The Landing Zone That Survived” — a year in the life of a New Zealand platform team, told from fakey.xyz. Fictional organisation, aggressively fake people, realistic problems.


Recap. The platform has onboarded two product lines, survived a 2am P1, taught a CFO’s four questions to answer themselves, and learned to build, route, or decline. Now the original mandate comes due: Barry Bigboss’ FY26 commitment to exit the Petone colocation facility — the rack of humming servers that justified this entire programme. This is the post about the work nobody puts on roadmaps: turning off the old world without breaking the new one. Plus: the annual plan lands, and a promise made in Part 6 comes due.


Here is the thing nobody puts in the landing zone reference architectures:

Building the platform is the easy half. The hard half is that everything the platform replaces has to actually stop.

The Petone colo had been fakey.xyz’s infrastructure home for eleven years. It held the legacy transaction store, the batch reporting estate, the DR target nobody had tested since 2022, and — the part that made this post necessary — the reason the hybrid network from Part 3 exists at all. The hub, the VPN, the corridor, the colo interconnects: all of it was architected around the assumption that Petone kept running. Exit the colo and the platform’s own topology has to change. The landing zone wasn’t just migrating workloads. It was migrating itself.

Robbie Deployment — who had decommissioned things before and carried the scars — put the framing on the wall of the war room on day one, and it’s the thesis of this post:

“Nobody celebrates a decommission. There’s no launch event for turning things off. Which is exactly why decommissions fail — everyone treats them as the boring tail of the project instead of the actual deliverable. The colo isn’t exited when the last workload moves. It’s exited when the last thing that depends on it — including us — is gone, verified, and the invoice is zero. Three different milestones. Two of them are where projects go to die.”


The Three Milestones Nobody Plans For

The war room board had three columns, and the discipline of the exit lived in keeping them separate:

Milestone 1: Workloads migrated. The visible, roadmap-friendly part. By the start of the exit window, only four workloads remained in Petone: the legacy transaction store (Vera Transaction’s long-tail archive), the batch reporting estate, the DR target, and — to everyone’s amusement — the building’s environmental monitoring system, discovered during the exit inventory running on a server nobody had logged into since 2021.

Milestone 2: Dependencies retired. Everything that referenced the colo without running in it: the VPN and its routing (Part 8’s incident had made this famous), DNS conditional forwarders pointing at colo resolvers, Serge’s log-shipping agent that had been quietly forwarding colo syslog for three years, monitoring probes, the corridor’s peering (which, with the transaction stream fully in NZ North, was now an exception to a requirement that no longer existed).

Milestone 3: Financial and contractual zero. The colo contract’s exit clauses, the hardware disposal with data-sanitisation certificates (a CPS 234 and Privacy Act evidence requirement — physical media with personal information doesn’t just “get thrown out”), the telecom circuits, the last invoice. Tessa Spreadsheet owned this column. Of course she did.

The exit inventory — a full dependency census before anything moved — was the single highest-leverage activity of the entire programme’s back half, and it found things no CMDB on earth would have:

  • The environmental monitoring system (already mentioned; it now runs on a $12/month Raspberry Pi in a corner of the Wellington office, and its migration is the most-loved entry in the platform’s migration log)
  • A nightly export job on the legacy transaction store that wrote to a colo-local NAS, which fed — nobody knew this — Tessa’s historical finance reconciliation. Tessa knew. Tessa had built it in 2019. Tessa had not told anyone. “It wasn’t shadow IT,” she said, with the exact cadence of someone describing shadow IT. “It was personal infrastructure.”
  • Serge Secure’s incident-response jump box, still living in the colo, referenced by an incident runbook from three CISOs ago. The platform’s reach matrix — built in Part 8 — flagged it, because the reach matrix’s whole design is dependencies as first-class data.

Lesson 1 — The exit inventory is a dependency census, not a server list. Runbooks, DNS, log shipping, monitoring, contracts, and one person’s 2019 side project all count. The census finds them; the CMDB won’t.


The Waves, Run in Reverse

The migration waves that onboarded workloads now ran in reverse, and the platform’s established machinery turned out to fit the exit almost eerily well — with one inversion nobody had planned:

The two-way contract, reversed. For each colo workload, the migration plan now had a “day after” section: what the workload team commits to post-cutover (their alerting, their cost dashboard, their onboarding into the workload-support stream) and what the platform commits to (connectivity, logging, the joint cutover checklist from Part 5 — which got its eighth and ninth uses, and its template updated each time).

The reach matrix earned its keep twice. Before each wave, the matrix answered “who depends on this colo component” — and twice it caught cloud-side dependents nobody had listed: a reporting job in Vera’s subscription reading from the legacy store, and a webhook from Netty’s scoring service to a colo-side enrichment API that had been scheduled for decommission but not yet executed. Both caught in planning, both trivially fixed in planning, both would have been 2am discoveries in execution. Fakey’s note: “The reach matrix cost one afternoon to build and paid for itself twice in one week. Build yours before your colo exit, not during.”

The Decommissioned management group, finally busy. The management group that had existed empty since Part 3 became the exit’s operational heart: every retired colo dependency landed there first — subscription moved, policies still applying, diagnostics still flowing, visibility retained — for a 30-day observation period before actual deletion. The policy: nothing is deleted directly from production. It is demoted to Decommissioned, watched for 30 days of silence (no access, no traffic, no alerts), and then deleted. The group caught two last-minute dependencies this way — a legacy auth callback and a billing job — both restored within hours from the demoted state instead of within days from backups.

Lesson 2 — Decommissioning is a monitored state, not an action. “Decommissioned” is a place with policies, logging, and a dwell time — not a verb.

The corridor, retired by its own register. The low-latency corridor — the exception that launched a thousand governance lessons in Parts 3 and 5 — reached the end of its natural life during the exit: the transaction stream it existed to reach now lived entirely in NZ North, so the exception had no reason to exist. And it was retired through the exception register, exactly as designed: quarterly review, evidence that the underlying requirement had expired, exception closed with documentation, corridor infrastructure demoted through Decommissioned. Serge Secure closed the register entry with what he later admitted was the closest he comes to sentiment: “First exception we’ve ever retired rather than renewed. Proof the register is a lifecycle, not a ledger.”

Netty’s farewell to the corridor, in the channel: “RIP to the bus lane. You carried my p99 for a year and never once asked me to slow down. This is more than I can say for most of my colleagues.”


The Week It Almost Went Wrong

Honesty requires the near-miss, because the near-miss is where the lesson lives.

The legacy transaction store — the last and largest workload — was scheduled for its cutover on a Wednesday, using the joint checklist, with rollback plans, with the colo contract’s exit date three weeks of buffer behind it. Comfortable. Almost too comfortable.

During the pre-cutover dependency sweep, Max Overhead — running the checklist’s platform section with the thoroughness of a man who had once been tested against it as a stranger — found the thing: Tessa’s 2019 NAS export job wasn’t the only reader of the legacy store. The store’s access logs showed a service principal, authenticating from inside the cloud platform, pulling data nightly — a principal created before the platform existed, owned by no current team, granted through RBAC inherited from a management group structure that predated everything in this series.

It took a day of archaeology to trace: a 2022 proof-of-concept’s data pipeline, built by a team that had since reorganised, still running on a schedule trigger in a logic app that everyone had forgotten, feeding a Power BI dataset that — this is the part that stopped the room — fed a dashboard the board had been seeing in every quarterly pack for two years. The board had been reading board reporting sourced from an orphaned PoC pipeline through a service principal nobody could name.

The response followed the series’ now-familiar shape: no blame (the pipeline’s original builder was traced, thanked for the documentation that made the trace possible, and invited to the retro), the dependency formalised (the board reporting is now fed by a sanctioned pipeline through the platform’s standard data pattern), the orphaned principal disabled through Decommissioned with the 30-day watch, and the exit checklist gained a permanent item:

“Access-log review of the workload being decommissioned: every principal, every path, last-used dates. An orphaned dependency is a decommission finding, a security finding, and a data-lineage finding simultaneously.”

Serge Secure’s evidence-pack entry, filed with visible satisfaction: “Exit process detected an orphaned service principal with production data access, created pre-platform, ownerless for 24 months. Found by checklist, remediated through the Decommissioned lifecycle, lineage restored. The exit demonstrated controls that didn’t exist when the principal was created.”

That’s the sentence worth sitting with: the colo exit found things the colo’s own security never had the vantage point to find — because the platform’s inventory, tagging, and logging gave the exit a completeness the old world had never possessed about itself.


Milestone Three: The Invoice Is Zero

Six weeks after the first wave, on a grey Wellington Thursday, Tessa Spreadsheet sent the shortest email of the series to the platform channel, Barry Bigboss cc’d:

“Petone colo: final invoice received, £0.00-equivalent. Contract terminated, hardware disposal certificates received and filed (sanitisation verified, CPS 234 evidence updated), circuit cancellations confirmed, environmental monitoring relocated to a Raspberry Pi that costs less per year than the old rack cost per hour. The exit is complete. Eleven years. Well done, everyone.”

Then, the line that mattered most, and the reason Tessa gets the closing credit of this arc:

“Spend next month will show the largest single reduction in the platform’s history, fully decomposed and attributed, in the standard reporting, with no manual mapping. As designed.”

The colo exit — the entire FY26 commitment — closed with the cost story legible to the board from the standard estate, because Parts 1 through 9 had built the machinery that makes a twelve-year-old datacentre’s death a two-line finance note instead of a three-week archaeology project.


The Annual Plan, and a Promise Kept

Two weeks later, the annual plan landed. Barry Bigboss kept the promise from Part 6 — and this is worth naming precisely, because keeping promises about platform funding is rarer than making them:

  • The platform funded to 2.4 FTE baseline — the gap Tessa had costed, closed in the plan, permanent.
  • The contractor requisition converted to a permanent role — Max Overhead, who had now run on-call through a real P1 and a colo exit, accepted. The rotation is three named humans. The Part 6 sentence — “the platform’s reliability is coupled to one person’s continued employment” — is formally retired, and Fakey read the retirement aloud at the team meeting, from the original email, for the drama it deserved.
  • The platform roadmap funded through FY27, with the BUILD items from Part 9 sequenced and the classification-aware diagnostics pattern resourced for generalisation.

Barry’s comment at the plan’s close, and Fakey’s record of it as the programme’s epitaph-in-progress:

“Twelve months ago I asked for a landing zone. What I got was a landing zone, an operating model, a cost story the board trusts, a security case the risk committee praises, and a datacentre exit delivered on the commitment that started it. The last one is what I’ll tell the board. The first one is what I asked for. The gap between what you ask for and what good platform work actually produces — that gap is why this team is funded properly now.”


The Increment Review Metrics

MetricTargetActual
Colo workloads migrated4 of 44 of 4 (+1 Raspberry Pi)
Exit inventory dependencies found pre-cutover11 (incl. 2 undocumented cloud-side readers, 1 orphaned principal, 1 environmental system)
Dependencies caught in planning vs at 2am100% pre100% pre
Decommissioned MG observation period honoured30 days30 days (2 dependencies recovered from it)
Exception register entries retired1 (the corridor, by lifecycle, documented)
Hardware disposal certificates (data sanitisation)filedfiled — CPS 234 evidence
Colo invoice$0$0
Orphaned principals with production data access found1, disabled, lineage restored
Platform FTE2.4 funded2.4 funded, permanent; rotation now 3 named humans

The Steal-This Checklist

  • Treat decommissioning as the deliverable, not the tail — three milestones: workloads moved, dependencies retired, invoice zero. Plan and staff all three.
  • Run a dependency census, not a server list — runbooks, DNS, log shipping, monitoring, contracts, board reports, and everyone’s personal infrastructure. The census finds what the CMDB can’t.
  • Decommissioned is a place, not a verb — a management group with policies, logging, and a 30-day silence watch before deletion. It will save you at least twice.
  • Build the reach matrix before the exit, not during — it catches cloud-side dependents in cheap planning time instead of expensive 2am time.
  • Review access logs of anything you’re turning off — every principal, every path, last-used dates. Orphaned access is a decommission, security, and lineage finding at once.
  • Retire exceptions through the register, by lifecycle — an exception that expires quietly is governance working; one that lingers forever is a liability with a review date nobody honours.
  • Get disposal certificates for physical media — data-sanitisation evidence is a compliance requirement, and it’s the part of cloud migration everyone forgets because the cloud made them forget hardware exists.
  • Fund the platform against its operating model, not its build — the build is a project; the operating model is forever. Budget accordingly, in writing, and keep the promise when the plan lands.

Next in the Series

Part 12 — “The Second Year Begins, or: The Platform Becomes Invisible.” The series finale: the annual review opens in recap, year two begins in vignettes, and Barry Bigboss asks the question the whole arc was building toward — what is a landing zone actually for?

One Block

Run a dependency census for one system you plan to retire: runbooks, DNS, monitoring, contracts, and personal infrastructure. Fix the first orphan you find.

What did your platform exit look like — and did anyone plan the dependencies, or did you discover them at 2am? Comments open at fakey.xyz.