Secure design principles
Pages
Page 5 of 17
3. Make disruption difficult
3.1 Ensure systems are resilient to both attack and failure
In order to cope with failure it is common practice to provide standby systems, alternative routes, and data backups. These perform well against random failure or mistakes, but often less well against malicious attack.
For example, if you have 10 identical load balanced servers and each has a 1 in 10 chance of random failure, the chances of them all failing at once are 1 in 10,000,000,000. However, if they all have the same vulnerability, it's very little extra work for an attacker to make all 10 fail rather than just one.
A second example is backups. Having a copy of your important data is a good idea. However, if an attacker can delete or corrupt the backups easily then they will only be useful to recover from random failures and mistakes.
Finally, for cyber-physical systems, safety controls are implemented to reduce the risk of a hazardous outcome. This process should carefully consider the possibility that an external threat actor could alter the integrity or availability of the safety controls.
Safety system architectures should be designed to ensure that unacceptable consequences are prohibitively costly for an attacker to achieve. To obtain this outcome, safety controls should be independent in the event of both a system compromise or a mechanical failure.
3.2 Design for scalability
To handle exceptional peaks in demand, or future expansion, you may need to scale your service quickly.
Considering the potential for future demand early on should mean your systems are easier to scale later. They should also be better suited to scaling out under increased demand, or when under attack.
3.3 Identify bottlenecks, test for high load and denial of service conditions
Identify any system bottlenecks. For example, low capacity, legacy business technology, or an essential microservice which calls a third party service. Ensure that you have a plan in place to handle these bottlenecks during periods of high load or outage.
Add specific tests for abnormally high load, and for denial of service, to your overall testing strategy. For instance, you could simulate some denial of service attacks by purposefully terminating certain microservices or infrastructure elements in your pre-production environments.
There are also openly available tools, such as Netflix's Chaos Monkey, which can help you test how your system will perform under high load or when components fail. It's important to test how you respond to failure conditions, as well as understanding what those failure conditions could be.
See also
NCSC guidance on defending Denial of Service attacks.
3.4 Identify where availability depends on a third party and plan for the failure of that third party
Many organisations rely upon third party services, such as telecommunication links, hosting, authentication or system administration services.
Ensure you understand the availability characteristics of these third party provisions and the impact on your operations should they fail, especially at times of critical demand. Have a plan for minimising disruption if such an event occurs.


