Who pulls the plug, what your cyber policy requires, which systems get preserved, and how long your last restore took

We shipped an incident response binder to our partners this week: roughly 2,000 pages, 49 playbooks, 11 incident families. A document that size is an odd thing to be proud of. It exists because the alternative is a four-page Word file from 2019 with a phone tree for two people who don't work there anymore, which is what I find most of the time when I ask to see an IR plan.

We built the binder because of OpenAI's letter last month on collective cyber defense, which 116 organizations signed on day one, ours included. Seth wrote about what it asks for a few weeks ago, so I'll skip the recap. The line that stuck with me was the one aimed at security companies, asking us to measure ourselves by how fast attacks get contained.

That's an odd thing to measure, because hardly any of it gets decided during the incident. It gets decided months earlier, when nothing is on fire and nothing feels urgent. I've sat in enough of these to know that the organizations moving fastest tend to be the ones where somebody already answered four questions while everyone was calm, which has less to do with tooling than people expect.

Those questions are below. None of them need the binder or a budget, but they do need preemptive consideration, so you’re not left scrambling when an incident does occur.

Who Has the Authority to Pull the Plug?

Ransomware spreads at the speed of your network, which means the window between noticing it and losing another few hundred machines is measured in minutes. Somebody has to decide to cut the connection, and that somebody has to be reachable, awake, and clear that the decision belongs to them.

In a lot of organizations, that person has never been named. The org chart doesn't settle it either, because seniority answers a different question than the one being asked. What you're really looking for is who is comfortable making that call and living with it either way.

Both mistakes are expensive in different ways. Shut down production without the authority and you've made a revenue decision that belonged to an executive, which you'll get to walk through in a meeting where you are the only agenda item. Wait on approval from someone who isn't picking up at 11 PM and the encryption keeps running while you hold, which costs more and is considerably harder to undo.

Organizations that have thought this through tend to define the line in advance. Some set a dollar threshold, above which an executive decides and below which the on-call engineer does. Some scope it by system, so cutting a branch office is a different conversation than cutting ERP. Some give the authority to whoever holds the pager and put in writing that the call won't be relitigated afterward. Each of those works. Improvising it at midnight is less effective.

Name the primary and the alternate, because incidents have a habit of starting while your CTO is somewhere over the Atlantic with the wifi off and a movie going. Then tell the people who might have to make the call that they're allowed to make it, which is the step that gets skipped. Authority that lives in a document nobody has read isn't going to help anyone at 2 AM.

What Does Your Insurance Require Before You Call Anyone?

Cyber policies carry conditions that read like procedural boilerplate right up until the moment you violate one. The common version requires notifying the carrier's breach hotline before you engage outside forensics, which means the incident response firm you've worked with for years might not be covered if you call them first. Some policies go further and restrict you to an approved panel entirely, so the question isn't only when you call, it's who you're permitted to call at all.

The reason this catches people is that the instinct in hour one is to get help on the phone, and the carrier feels like paperwork that can wait. I've watched a team work that out at hour six, on a conference bridge, while somebody reads coverage language aloud to a room that has other problems. The answer they got was expensive.

The fix costs one afternoon. Pull the insurance policy and find three things: the notification requirement and its deadline, whether your preferred IR firm is on the approved panel, and what the carrier expects you to preserve before anyone starts remediating. Then put the hotline number on the same sheet as the rest of it and make sure the person most likely to pick up the phone at 2 AM has seen that sheet.

Which Systems Get Preserved Before Anything Gets Rebuilt?

Skip the preparation and people improvise. It usually goes one of three ways.

Somebody reimages the machine that held the evidence, because the machine looked infected and rebuilding it felt like progress. Somebody resets the password on the account the attacker is sitting in, which tells the attacker they've been spotted and prompts them to move faster or start deleting. Somebody sends the breach notification from the tenant that's still compromised, giving the attacker a copy of the legal strategy.

Each of those is a reasonable instinct executed at the wrong moment by a person who had thirty seconds to think about it, and each one costs days on the back end.

Two decisions made in advance prevent most of that. The first is which systems get preserved before anything gets rebuilt, which usually means agreeing that the first compromised host gets imaged and left alone even when everyone wants it back in production. The second is who confirms an account is clean before anybody touches the credentials, so the containment work happens in an order that doesn't announce itself.

Those conversations take twenty minutes. They're also uncomfortable enough that they slide to next quarter, every quarter.

When Did Someone Last Restore From Backup, and How Long Did It Take?

Plenty of organizations can confirm the backups are running. The monitoring dashboard says green and the job completed last night. Fewer can name the person who has performed a full restore recently enough to do it under pressure. Fewer still can tell you how many hours it takes, which is the number that determines whether you're down for an afternoon or a week.

A related question matters just as much: whether your backup administrator and your domain administrator are the same account. When they are, an attacker who takes one has taken both, and the recovery plan you were counting on becomes a negotiation instead. Ransomware crews look for backup infrastructure early and deliberately, because a victim who can restore is a victim who doesn't pay.

What settles this is a restore test you schedule, with a stopwatch running and a person who isn't the usual person doing it. Time it, write the number down, and check whether the credentials that reach your backups reach anything else.

If the answer is yes, separating them is a weekend of work that buys back your leverage.

What to Do With the Next Afternoon You Have

The OpenAI letter asks organizations to fix their highest-risk weaknesses and to build in least privilege, strong access controls, and defense in depth, which is advice that predates it by roughly two decades. It would be easy to read least privilege, access controls, and defense in depth as a weak response, decades-old advice attached to a threat that's brand new. I'd read it the other way. If the guidance had changed, everyone would be starting over. It hasn't, which means the work in front of you is work you already know how to do, and none of it is waiting on the next model release or anyone else's commitments.

The four conversations are the same kind of thing. An afternoon each, no SOC required, no IR retainer, no line item, which matters more than it sounds, because a good number of the organizations reading this are running on a small IT team, an outside provider, or one very tired person who is good with computers. Those are the environments where a decision made in advance matters most, because there's nobody sitting idle to figure it out live.

If you want an order, I'd go authority first, since it's the shortest conversation and it unblocks every other decision in the first hour. Insurance second, because the answer might change who you're allowed to call. Evidence preservation third. Backups last, only because that one turns into a project rather than a conversation, and projects lose to conversations when you're trying to build momentum.

The binder helps with the rest of it, though a binder on a shelf is a doorstop, albeit a sturdy one. Pick the three incident families you're most likely to face, read those playbooks while you're bored, fill in the parts we can't possibly know, and run a tabletop with leadership in the room instead of just IT. A tabletop with only technical people in it tests the technical response, which was probably fine already. The decisions that slow organizations down get made above that line.

So here's the question I keep asking people: if ransom notes showed up on every screen tonight, who has the authority to take the network offline, and do they know they have it? Answer without hesitating and you're ahead of a lot of organizations your size. Hesitate, and you've found where to start on Monday.

The IR playbook binder is available to all Galactic partners. If you're interested in becoming one, let's chat. Schedule some time here.