Skip to main content
Annual Review

NCSC Annual Review 2025

Looking back at the National Cyber Security Centre's ninth year and its key developments and highlights, between 1 September 2024 and 31 August 2025.

Page 16 of 32

Engineering resilience against critical loss

Prevention and detection aren’t enough: resilience means building systems which can operate and recover following a disruptive cyber intrusion.

The cost to Marks and Spencer and its insurers following the ransomware attacks earlier this year is estimated to exceed £300m. However, the impacts of cyber attacks are now felt beyond the victim organisation. Destructive cyber intrusions increasingly affect customers, third parties, and wider society.

For example, the ransomware incident suffered by pathology laboratory services provider Synnovis led to significant clinical healthcare disruption across the London region, and incurred costs of £32.7m, far outstripping Synnovis’ profits of £4.3m for 2023. It also directly contributed to at least one patient death. This starkly illustrates the scale of exposure across our hyper-connected and technology-dependant society, especially where legacy infrastructure and systems need to be maintained.

These challenges are only magnified when considering the technical legacy often held in organisations. Having a strategy to address technical legacy is a foundational precursor to engineering resilience. Recognising this, the NCSC included an emphasis on understanding threat actor methods and motivations coupled with monitoring and threat hunting in the Cyber Assessment Framework v4.0 (Cyber Assessment Framework v4.0 released in response to growing threat ).


How can resilience engineering be applied to the cyber security domain?

In designing systems to be resilient we can draw on several existing and emerging cyber security architectural and operational approaches.

Infrastructure as code allows us to rapidly and reliably replicate and reconstitute systems and infrastructure. This is an essential component of rapid recovery from loss or compromise, but also allows the deployment of trusted, immutable infrastructure making it far more difficult for threat actors to maintain persistence.

Similarly having evidence of immutability of backups and that there are practised recovery and rehydration procedures for situations where total environment loss, including the loss of identity and access management, hypervisors, cloud configurations and more, has occurred is essential for timely and effective recovery.

Segmentation through both logical and physical architectural patterns allows for isolation and contained operation either as-needed to minimise impact during an event, or persistently to create trust boundaries such as through cross-domain solutions. The NCSC has emergent anecdotal evidence that those organisations who intervene during a destructive event and self-isolate, recover quicker with less impact. Approaches including Privileged access Workstations (PAWS) and segregation of management planes enhance resilience against systems administrator compromise by materially complicating total environment compromise and loss through privileged access.

Applying the principle of least privilege universally and across all services further limits the potential damage from breaches. This approach should be complemented by the use of microservices, context-aware and individualised authorisation, and application isolation, all of which reduce the ‘blast radius’ of a compromised component or service. These principles underpin a Zero Trust Architecture (ZTA), which assumes no implicit trust within the system.

Observability and monitoring are key enablers for Engineering Resilience, to enable observation and rapid response to events. Comprehensive observability and data collection (hosts, infrastructure, edge, cloud, and applications etc.) provides opportunities for anomaly detection, response and post incident learning, which are critical for minimising impact and improving future resilience.

Chaos engineering, the deliberate introduction of failure to validate detection and recovery, is a complementary approach to evidencing control efficacy in dealing with systems failure and the ability to self-heal. Some contemporary organisations repeatedly test through such approaches that resilience exists against partial and sudden loss along with the ability to detect and respond.

In the most critical situations having critical business functions running on duplicate but distinct technology stacks can provide resilience against a range of operational and cyber security related risks.


Published

Reviewed

Version

1.0