Part 6 of “The Landing Zone That Survived” — a year in the life of a New Zealand platform team, told from fakey.xyz. Fictional organisation, aggressively fake people, realistic problems.
Recap. Increment Two put a production workload — Netty Latency’s fraud scoring service — on the platform, survived an eleven-minute DNS incident and a 2am quota surprise, created a paging path that is formally “Robbie, primary; Fakey, secondary; both of us hoping the bus factor holds until hiring”, and saw Fakey McFakerson tell Barry Bigboss in writing that the platform’s reliability is coupled to one person’s continued employment. This is the post where that sentence comes due.
Every platform has an operating model. Most organisations just haven’t written theirs down, so it exists only as folklore — who actually gets called when things break, and who actually answers.
At fakey.xyz, the folklore version went something like this: alerts go to the platform channel; Robbie Deployment notices them; Robbie Deployment fixes them; Robbie Deployment is a person, and this arrangement scales exactly as well as that fact.
Increment Three was scheduled to be the expansion increment — second product line onboarded, sandbox automation, vending version one with Robbie’s “future pipeline logic” log becoming code. Instead, it became the operating model increment, because of a Thursday, and a near-miss, and a spreadsheet that Tessa Spreadsheet brought to a meeting she wasn’t invited to.
The Thursday That Changed the Agenda
The near-miss was almost embarrassingly small.
A cost alert fired. One of the two enterprise clients’ migrations had a misconfigured snapshot policy that was quietly accumulating storage at a rate that would have cost fakey.xyz roughly $4,200/month by quarter end. The alert went to the cost alerts email distribution list.
That distribution list — created in increment one, in five minutes, by a team that had bigger problems — contained: Fakey (correct), Tessa Spreadsheet (correct), Robbie (correct), and a product manager who left fakey.xyz four months earlier (less correct). What it did not contain was anyone whose job included triaging cost alerts. Tessa saw it a week later during her monthly cost review, forwarded it to the platform channel with the message:
“This alert is 9 days old. It is correct. Who was supposed to action it? Asking for finance’s entire confidence in this platform.”
Nine days. The alert had worked perfectly. The operating model had done nothing, because the operating model — when you actually traced it on a whiteboard, which is what Fakey did that Thursday afternoon — looked like this:
- Cost alerts → email list → triaged by whoever’s monthly ritual happened to surface them (Tessa)
- Security/detection alerts → Serge Secure and Locksley Keymaster, 20% allocated, whose alert triage competed with their actual day jobs
- Platform health alerts → Robbie, who also carried the on-call pager, the vending pipeline, the migration waves, and a backlog
- After-hours platform P1s → the pager → Robbie
- Robbie’s leave, illness, or resignation → nothing, formally; Fakey, informally
Four alert streams, four different destinations, zero shared triage, no SLAs, no ownership boundaries between platform obligations and workload obligations, and a single human being load-bearing at the centre of all of it.
Fakey took a photo of the whiteboard and sent it to her own phone with the caption “the actual architecture.”
The Meeting Tessa Wasn’t Invited To (And Invited Herself To)
The following week, Fakey convened the operating model working session: herself, Robbie, Serge Secure, Locksley Keymaster, and — she assumed this was uncontroversial — a plan to keep finance out of it, because the operating model was an engineering conversation.
Tessa Spreadsheet, upon learning of the session’s existence, invited herself, arrived with a spreadsheet, and opened with:
“You’re about to design who watches the alerts. Every alert you watch costs either headcount or attention, and both come out of the same envelope I report on. Also, ‘cost alerts have no owner’ is in my opinion the single worst line on your whiteboard, and I don’t say that because it’s mine — I say it because a cost alert that nobody actions is exactly as dangerous as a security alert nobody actions, and Serge would never tolerate his being unowned. So why is mine?”
Nobody had an answer. Tessa stayed for the whole session. By the end of the day she was, formally, the owner of the platform’s FinOps operations stream — and the operating model had gained its first and most important property: it was written down, with names, in a document that finance and security and engineering had all edited.
The Operating Model, as Designed
What fakey.xyz landed on — after that session, two more, and one uncomfortable conversation with Barry Bigboss that we’ll get to — is presented here in the shape you can steal. Five streams, one document, one principle.
The principle first
Every alert, every recurring task, and every failure mode has a named human owner, a response commitment, and a way to escalate that doesn’t depend on who happens to be reading the channel.
Not a team. A named human. Teams don’t triage alerts; nervous systems attached to job descriptions do. The operating model document has a column called “who, by name” and Fakey has personally rejected two drafts that wrote “the platform team” in it.
The five streams
| Stream | What it covers | Owner (named) | Commitment | Escalation |
|---|---|---|---|---|
| Platform health | Hub, firewall, workspace, vending, DNS, identity | Robbie (primary), Fakey (secondary) | P1: page, 15-min ack. P2/P3: channel, next business day | Fakey → Barry |
| Security & detection | Sentinel/detection rules, identity anomalies, policy drift | Locksley Keymaster (primary), Serge (escalation) | P1: page, 15-min ack; others daily triage | Serge → Risk committee |
| FinOps | Budget breaches, anomalies, orphaned resources, allocation | Tessa Spreadsheet (primary), monthly review with Fakey | Anomaly: 2 business days. Budget breach: same day notify | Tessa → Penny Pockets |
| Workload support | “How do I”, access requests, environment questions | Platform channel, owned by Fakey, rostered | Same or next business day | To platform health if it’s an outage |
| Compliance evidence | CPS 234 pack, exception register reviews, Privacy Act tabletops | Serge (quarterly), Locksley (continuous collection) | Quarterly cycle, calendar-blocked | Serge → Board risk committee |
The boundaries — the part everyone gets wrong
The hard-fought part of the model wasn’t the table. It was three boundary decisions:
1. Platform obligations vs workload obligations, written as a contract. The platform owns the hub, the firewall, the workspace, the vending pipeline, the policy estate, and DNS. The workload owns everything inside their subscription: their app, their app’s alerting, their data, their recovery. This was already informally true — the Friday incident retro had proven the value of the “platform broke vs workload coped” distinction — but now it’s a document, linked from every vending output, that workload teams explicitly acknowledge when they receive their environment. Netty Latency reviewed it and suggested one change that made it dramatically better: a column listing what the workload team can expect the platform to do during a platform-caused incident (communicate, restore, provide a post-incident report within five business days). Promises flow both ways or they aren’t promises.
2. Alert fatigue managed by budget, not by heroics. Every alert rule in the monitoring estate must justify its existence against the streams table: if it fires, who acts, and what do they do? Serge’s estate and Robbie’s estate were audited against this in one afternoon — 14 alert rules were deleted or demoted to informational because their honest answer to “who acts on this?” was “nobody, it’s just interesting.” Locksley Keymaster’s summary: “An alert nobody actions is not monitoring. It’s ambient anxiety.”
3. The after-hours question, answered honestly for 1.6 FTE. Here is the decision Fakey describes as the one she is proudest of in the entire programme, and it was only possible because she was willing to look cheap in front of her CEO:
The platform’s formal after-hours commitment is P1-pageable only, business-hours-adjacent, with a documented restoration target of 4 hours, and — critically — no production workload is permitted to onboard a hard real-time dependency on the platform’s after-hours response without an explicit, written risk acceptance from their GM.
Netty’s team accepted that risk for the fraud scoring service — because their circuit-breaker architecture (the one that saved Friday) means platform outages degrade them rather than kill them, and they said so in writing. That letter — a workload team documenting their own resilience as a condition of accepting a platform limitation — is the operating model working exactly as designed. Honesty about limitations, it turns out, produces better customer architectures than pretending to capabilities you don’t have.
The Barry Conversation, or: The Cost of an Operating Model
Which brings us to the uncomfortable conversation. Because every one of those named owners is a person with a finite week, and the streams table — audited honestly — consumed capacity that didn’t exist.
Tessa, who had by now fully assumed the role of the platform’s most dangerous ally, prepared the analysis. Her spreadsheet modelled the platform team’s actual committed hours: product delivery work that couldn’t be dropped, the migration waves, the increment roadmap — and the operating model’s ongoing obligations: on-call, triage, the fortnightly coffee follow-ups from Discovery Week, the cross-team reviews from the DNS retro, quarterly compliance cycles. The number that fell out, rounded with cruelty:
The platform, as promised, requires approximately 2.4 FTE. The platform, as funded, has 1.6 FTE. The gap is being paid, right now, in Robbie Deployment’s evenings, Fakey’s fractional-product-role erosion, and alert streams that are one resignation away from silence.
That one-page analysis — Tessa called it “the true cost of the platform, versus the price we’re currently not paying” — went to Barry Bigboss with Fakey’s covering note, which did not ask for sympathy. It asked for a choice between three named options:
- Fund to 2.4 FTE at the annual plan (five months away), with the gap bridged by pausing the second migration wave. Cost: schedule. Risk: low.
- Bridge immediately with a 6-month contractor to carry on-call and workload support. Cost: money now, from an envelope that didn’t exist — which made this partly Penny Pockets’ decision. Risk: knowledge transfer overhead.
- Continue at 1.6 FTE and reduce the promise — shrink the after-hours commitment, defer vending automation, slow onboarding. Cost: the platform’s momentum and credibility, and Barry’s own FY26 datacentre-exit commitment. Risk: the shadow-IT dynamics from the very first series of posts, returning.
Fakey’s recommendation was option 1, with option 2 as the bridge if the annual plan couldn’t be pulled forward. And then she did the thing this series has been building toward since Part 1 — she made the consequence concrete and personal:
“Barry, option 3 means the platform’s reliability remains coupled to one person’s continued employment. You have that in writing from me since last increment. Robbie Deployment is currently the platform. If he leaves before the annual plan, option 3 becomes unavailable and we will be choosing between 1 and 2 in an emergency, at emergency prices, with an emergency timeline.”
Barry’s answer, delivered in the meeting: “Pull the contractor decision to the next exec meeting. I’ll bring Penny. And Fakey — when the annual plan lands, this platform is funded properly. I’d rather explain a bigger platform line to the board than explain a payments outage.”
Penny Pockets approved the six-month contractor bridge the following week. The requisition went out for a platform engineer — senior, NZ-based, Terraform-native. The hiring market being what it is, this will take weeks, and the series will feel it.
But the operating model decision was made, and it was made before the emergency, with options, a recommendation, and a named cost — which is the entire difference between governance and improvisation.
The Increment Three Review, Reframed
Increment Three shipped less technology than planned — sandbox automation slipped, vending version one slipped, the second product line’s onboarding started but didn’t finish. The metrics table told that story honestly:
| Metric | Target | Actual |
|---|---|---|
| Second product line onboarded | complete | started (2 of 4 environments live) |
| Vending version one (automated) | shipped | slipped — “future pipeline logic” log 80% complete |
| Alert streams with named owners and SLAs | — | 5 of 5 |
| Alerts deleted as unactionable | — | 14 |
| Cost alerts triaged within commitment | — | 100% (Tessa’s stream, since day one) |
| Platform FTE gap documented and escalated | — | documented, escalated, bridged (contractor approved) |
| Robbie Deployment’s after-hours incidents | — | 2 (both under 30 min, both post-incident-reviewed) |
Barry Bigboss — and this is the moment the series has earned — looked at the slipped rows, looked at the operating model rows, and said: “You shipped less platform and more organisation. That’s the right trade. Say that in the next board update, in those words.”
The Steal-This Checklist
- Trace your actual operating model on a whiteboard — every alert stream, every recurring task, every escalation path, as it really runs. Photograph it. Caption it honestly.
- Five streams, named humans, written commitments — platform health, security, FinOps, workload support, compliance evidence. “The platform team” is not a name.
- Finance owns cost alerts, formally — an unactioned cost alert is exactly as dangerous as an unactioned security one, and finance makes the best FinOps owner you’ll ever get.
- Audit your alert estate against “who acts on this?” — delete or demote what fails. Ambient anxiety is not monitoring.
- Platform vs workload obligations as a two-way contract — including what the platform promises during its own incidents.
- Honest after-hours commitments for your actual team size — and let workload teams accept the residual risk in writing. Their resilience architectures are better than your pretences.
- The true-cost analysis: promised capability vs funded capacity, in FTE, in one page — with three named options and a recommendation, presented before the emergency, not during it.
- Escalate the bus factor in writing, repeatedly, until someone with budget acts on it.
Next in the Series
Part 7 — “The Cost Surprise, or: The CFO’s Four Questions.” The contractor starts, the migration waves accelerate — and cloud spend doubles in a quarter. Penny Pockets summons the platform team with four questions about the board report, and the platform discovers whether the subscription strategy, the tagging estate, and Tessa’s allocation rule were the insurance policy they were designed to be. Also: the workstation under the desk finally makes its move.
One Block · build from here
What does your platform’s operating model actually say — and is there a version of it written down that would survive the person who currently holds it together leaving? Comments open at fakey.xyz.