Ask a founder what happens if a server starts encrypting itself at 2 a.m. and the honest answer, most of the time, is a shrug: "we'd figure it out." That sentence is the plan. It names no one, decides nothing in advance, and assumes that a group of tired people under financial and reputational pressure will make good, fast, well-sequenced decisions about systems they have never had to reason about this way before. Sometimes that goes fine. Often it does not, and the difference rarely comes down to the sophistication of the attack — it comes down to whether anyone had already thought through what happens next.
That is the part that is fixable in advance, cheaply, before anything is wrong. This piece is about what "before" actually consists of.
Why improvising during the incident fails
The failure mode is not usually a lack of technical skill. Most IT teams can patch, rebuild, and restore from backup. What breaks down is everything upstream of that: who decides, how fast, and on what information.
Decisions get made under pressure by whoever is in the room. A breach rarely announces itself as a breach. It shows up as a slow file share, a strange login alert, or a call from a customer asking why their invoice looks wrong. Someone has to decide, in real time and usually alone, whether this is worth escalating, whether to take a system offline, and whether to tell anyone outside the immediate team. Without a predefined threshold for those calls, the person closest to the alert is making a judgment call that should belong to several people, and they are making it fast, with incomplete information, and often without the authority to actually pull the trigger.
Well-meaning cleanup destroys the evidence needed to understand what happened. The instinctive response to a compromised machine is to fix it: reboot it, reimage it, delete the suspicious files, restore from a backup and move on. Every one of those actions can erase exactly the evidence an investigation needs. A reboot clears the memory that would have shown which process was actively communicating with an attacker. Reimaging destroys the disk artifacts that would explain how the attacker got in. Deleting the ransom note or the suspicious binary removes the sample that would identify the variant and whether a decryption weakness is known. None of this is negligence — it is the normal, sensible instinct to stop the bleeding, applied by someone who has never been told that stopping the bleeding and preserving the wound are sometimes in tension.
Nobody has the standing authority to take systems offline. This is the quietest failure and the most common one. Disconnecting a production system — even briefly — has revenue, customer, and reputational consequences, and in most companies no single person below the executive team feels entitled to make that call unilaterally. So it gets escalated, and escalated again, and by the time someone senior enough says yes, the attacker has had hours of uninterrupted access. A plan that names who can authorize containment, and under what conditions they don't need to ask first, removes that entire delay.
The common thread across all three is that they are not technology problems. They are organizational and procedural gaps, and organizational gaps are exactly the kind of thing you can close on a quiet Tuesday afternoon rather than during an active compromise.
What readiness actually consists of
"Have an incident response plan" is true but not specific enough to act on. In practice, readiness for a company without a dedicated security team comes down to four concrete things, and each is worth building separately rather than treating "readiness" as one large, deferred project.
- A written incident response plan and runbooks. Not a philosophy document — a short, specific set of steps for the scenarios you can reasonably anticipate (ransomware on a file share, a compromised admin account, a suspicious login from an unfamiliar country, a possible business email compromise). Each runbook states the first three or four actions, in order, and who is authorized to take them without further sign-off.
- Defined roles and an escalation path. Who is the incident commander? Who can authorize disconnecting a system from the network? Who talks to customers, regulators, or the press, and who explicitly does not? This can be written down in an afternoon and it removes the single most common source of delay: nobody knowing whose call it is.
- Logging and evidence-preservation basics, in place before anything happens. An investigation is only as good as the record it can examine. Centralized logs with a stated retention period, authentication and access logs that are not immediately overwritten, and endpoint visibility on critical systems are what turn "we think something happened" into "here is what happened, when, and how." If this is switched on the week of the incident, there is nothing from before that week to examine — the most useful evidence is always the evidence collected on ordinary days when nothing looked wrong yet. Reducing the number of exploitable gaps in the first place, through routine penetration testing rather than an automated scan alone, lowers the odds you need any of this — though it is a mitigation, not a substitute for readiness, since testing reduces likelihood and readiness reduces damage once something gets through anyway.
- A pre-agreed relationship with a forensics and incident response provider, ideally a retainer or at minimum a known point of contact, agreed before you need it. The alternative — searching for "incident response near me" while systems are actively encrypting — costs you the hours where speed matters most, and it means the first conversation with whoever you hire is also the first time they have ever heard of your environment. A digital forensics and incident response retainer exists precisely to remove that search from the worst possible moment to be doing it.
None of this requires a security operations center or a large budget. It requires deciding these things once, writing them down, and revisiting them occasionally — closer to a fire drill than to a standing department.
Forensics and response are related but distinct disciplines
The two terms get used almost interchangeably, and the distinction is worth being precise about, because they answer different questions and often run on different timelines.
Digital forensics establishes what happened: how the attacker got in, what they accessed, how long they were present, and what evidence supports each of those claims — under a documented chain of custody, because the findings may need to hold up to an insurer, a regulator, or a court. Incident response is the operational work of containing the incident and recovering safely — isolating affected systems, closing the entry point, and restoring service without reopening the same door. In practice the two run together: containment decisions are made with an eye to what evidence they will preserve or destroy, and the investigation shapes what "safe to restore" actually means.
A useful test: if the question is "is this thing still doing damage right now, and how do we make it stop," that's response. If the question is "what exactly did they touch, and can we prove it," that's forensics. Most real incidents need both, run by people coordinating closely rather than in sequence, which is the whole argument for treating digital forensics and incident response as one engagement rather than two separate purchases made under pressure.
The first hour, done properly
It's worth walking through what a properly prepared first hour looks like, because it is the clearest illustration of why the pieces above matter together rather than individually.
Someone notices the signal — a spike in failed logins, an EDR alert, a report from an employee that their files look wrong. Because roles were defined in advance, they know who to tell, and that person has the standing authority to isolate the affected system from the network without first convening a meeting. Because the runbook says so, the system is disconnected but not powered off, preserving the memory contents that would otherwise be lost. Because logging was already centralized, there is a record of what happened in the hours and days before anyone noticed, not just from the moment of detection onward. And because a forensics and incident response relationship was already in place, the call to bring in outside help happens in the first hour rather than the third day, and it goes to a provider who can move immediately instead of one who has to be briefed on your environment from zero. None of that requires improvisation, because none of it was left to be figured out live — and the same visibility that supports an investigation, built through a network security assessment that maps what's actually on the network and how traffic moves across it, is what allows an investigator to trace how someone moved once they were in and to state with the record where the boundaries actually sat.
That contrast — decisions pre-made versus decisions improvised — is the whole difference between a contained incident with a defensible account of what happened and a longer, more expensive, less certain version of the same event. The technical difficulty of the breach itself is often similar in both cases. What differs is how much of the response was already decided.
Where this lands
Readiness is not a large project and it does not require assuming the worst about your own security. It requires writing down, once, who decides what, making sure there is something to investigate if it ever comes to that, and knowing who you would call before you need to call anyone. Most of that work happens in a single focused engagement rather than an ongoing program, and it is far cheaper and calmer to do on a normal week than to attempt for the first time during an active compromise. If you are trying to work out how much of this you already have in place and where the actual gaps are, that is a straightforward scoping conversation to have with us, and it usually starts with a short review of what you're logging and who currently has the authority to pull a plug.