Englishالعربية Soon
Under attack?
Root cause analysis

Removing the malware is not the fix. It is the cleanup.

A structured investigation into why an incident was possible and why nothing stopped it earlier, ending in corrective actions with owners, not a report that says “improve security awareness”.

When to commission one
After a confirmed breach or ransomware event
After a near miss that nobody can explain
When the same class of incident keeps recurring
When a regulator or insurer asks for one
After a failed audit or control breakdown

An RCA is blameless by design. The question is never who clicked the link. It is why clicking a link was sufficient to compromise the estate.

The method

Keep asking why until the answer is something you can assign

An illustrative chain from a real pattern we see often. Each step down is a level of causation, and only the last one produces a corrective action worth funding.

Symptom
The finance workstation was encrypted by ransomware.

This is where most incident reports stop, and why the same thing happens again nine months later.

Why 1
Because an operator executed a malicious attachment.

True, and useless on its own. People will always eventually open something.

Why 2
Because the attachment reached the mailbox and was executable.

Gateway policy permitted the file type, and no detonation was performed before delivery.

Why 3
Because the endpoint had no behavioural detection in enforcement mode.

The agent was deployed in audit-only mode after a false positive eighteen months earlier, and never switched back.

Why 4
Because the exception was never reviewed and nobody owned it.

Temporary exceptions were recorded in a ticket, not in a register with an expiry date and an owner.

Root cause
Because there is no process for expiring security exceptions.

This is fixable, assignable, and prevents an entire class of future incidents rather than one.

In practice the chain is rarely linear. Most incidents have two or three causal branches, and the analysis is only finished when each branch reaches something structural.

Techniques

Six techniques, selected to fit the incident

We do not apply all of them to everything. The choice is made once the timeline exists, because that is when the shape of the failure becomes clear.

01

Five whys

A disciplined descent from symptom to structural cause. Fast, and enough for most single-thread incidents.

02

Fishbone analysis

Contributing factors sorted by category: people, process, technology, governance, when the cause is not a single chain.

03

Fault tree analysis

Working backwards from the failure through the conditions that had to combine for it to occur. Used where the impact was severe.

04

Timeline reconstruction

Every logged and reported event placed in sequence, which is how gaps between what happened and what was noticed become visible.

05

Control failure mapping

Each control that should have prevented, detected or limited the incident, and the specific reason it did not.

06

Barrier analysis

What stood between the threat and the asset, which barriers were missing, and which were present but ineffective.

Contributing factors

Four categories, because the cause is almost never only technical

Every incident is examined against all four. An analysis that finds only technology causes has usually stopped early.

People

Skills, staffing levels, workload, awareness, and whether the person who noticed knew who to tell.

Process

Change control, exception handling, patch cycles, escalation paths and whether the documented process is the one in use.

Technology

Coverage gaps, tools in audit-only mode, missing telemetry, unmanaged assets and end-of-life systems.

Governance

Ownership, risk acceptance, budget decisions and whether anyone was accountable for the control that failed.

Engagement

Three weeks, then a check at ninety days

Days 1 to 2

Evidence and scope

Logs, forensic artefacts, tickets, change records and configuration state collected while they still exist. Scope and the questions to be answered agreed with you in writing.

Days 3 to 5

Timeline

Every event placed in sequence: what the attacker did, what was logged, what alerted, what a human saw and when anyone acted. The gaps between those four lines are the finding.

Week 2

Interviews

Blameless conversations with the people involved. Almost every analysis turns on something that was known informally and never recorded anywhere.

Week 2 to 3

Causal analysis

Techniques applied, causal chains followed to structural causes, and each failed control mapped to the specific reason it did not work.

Week 3

Report and actions

Findings, causal chains, evidence, and a corrective action plan with owners, effort and dates. Presented to your technical team and separately to leadership.

Day 90

Verification

We come back and check whether the corrective actions were implemented and whether they hold. This is the step most organisations skip and later regret.

Deliverable

What you receive

A written analysis with the causal chains, the evidence behind each one, the control failures mapped individually, and a corrective action plan stating owner, effort, dependency and date for every item. Plus a short version for the board and, where required, a factual account suitable for a regulator or insurer.

Incident response →

Deliverable

What we will not write

“Human error” as a root cause. “Improve security awareness” as a corrective action. A finding that names an individual. Any conclusion the evidence does not support, if the logs were not retained, the report says so rather than guessing.

Talk to us →

Related

What sits either side of an analysis

Incident response

Containment first, analysis second. The evidence an RCA needs is preserved during response or not at all.

Compromise assessment

If you suspect something happened but have no confirmed incident, start by establishing whether you are compromised.

Risk assessment

Corrective actions belong in the risk register, or they will be forgotten by the next budget cycle.

Find out why it was possible.

If the incident was in the last ninety days, most of the evidence still exists. After that, the analysis gets harder and thinner.