Building a Security Operations Centre (SOC)
Pages
Page 12 of 14
Detection practices
Some things you should consider when you are building a detection capability in addition to the techniques already covered.
Baseline comparison
It is common to perform a baselining exercise when building a SOC. The aim of this exercise is to ascertain exactly what ‘normal’ looks like within an organisation. In theory, this is a valid approach, but you should be wary of making assumptions. This is a hard task, especially in large organisations.
Consider the following questions:
How do I know that this system isn't already compromised?
- If it is, your baseline will include any malicious traffic and you may never detect the attack.
What happens with new systems and developers?
- Change control is an important part of this, but the reality is, not everything is caught in that process. Be sure you understand the cost of triaging new systems and development activities.
Note:
Baseline comparison is also one of the key techniques used to detect malicious behaviour within legitimate use of a system. By profiling what normal user behaviour looks like it is possible to detect usage that could be malicious or even fraudulent. This is covered in more detail in our Transaction Monitoring guidance.
Single pane of glass
Indications of an attack will rarely be isolated events on a single system component or system. So, where possible, having a single platform where analysts have the ability to see and query log data from all of your onboarded systems is invaluable.
Having access to the log data from multiple (or all) components, will enable analysts to look for evidence of attack across an estate and create detection use-cases that utilise a multitude of sources.
By creating temporal (actions over a period of time) and spatial (actions across the estate) use-cases, an organisation is better prepared to address cyber security attacks that occur system wide.
Triage as an objective
When creating an alert, spare a thought for the analyst who will triage it. The more information that an alert pulls from available log sources, the quicker it can be investigated, or dismissed as a false positive. Often organisations will implement 'runbooks' that allow analysts to react promptly to an alert. This is a contentious subject as each alert should merit its own investigation from a capable security analyst. Whereas, following a specific script within a runbook, might limit the level of genuine investigation. Conversely, guidance on how to triage alerts, detailing which log sources to examine, which systems to investigate and who to contact are of course, extremely useful.
This is why SOCs will always try and ensure that there are the absolute minimum number of alerts, so that when an alert does fire, it can be properly investigated and the SOC can react accordingly.
False positives
Working out how to effectively and safely address false positives is a challenge for all SOCs. However, there are a few common approaches that attempt to reduce false positives.
- Triage feedback - Building a feedback process into the triage capability of your SOC will enable you to refine the logic used within a detection use-case.
- Allow listing - Common for some development environments, where change is frequent. Specific users or IP addresses can be excluded from some (specific) use-cases until a more stable time. It is very important that this is strongly object/time bounded, to minimise the risk of perpetual allow listing and the evasion of monitoring.
- Combining alerts - As real attacks will often trigger multiple detections, when a detection alert fires in isolation, it is more likely to be a false positive. This clearly depends on the logic of the detection. By implementing alerts based on real attack profiles, i.e. multiple detection use-cases, you can attempt to minimize false positives. This is especially prudent when dealing with more sophisticated actors, they may not trigger single alerts - success can be had by combining multiple lower severity alerts, however this needs to be balanced to ensure that single alerts aren’t ignored.
Use-case lifecycle
You should not allow the rulesets which drive your detection process to stagnate, because the threat landscape, the organisation and the technology it uses will all change over time and these changes need to be matched by your use-case development.
A good approach to the problem is to implement a use-case lifecycle, to ensure that the use-case remains valid throughout its use. This should also include regular reviews of commonly triggered alerts, to ensure that these alerts are still useful to the SOC on balance and are not swamping analysts with false positives.
Reviews should look for opportunities to improve rule logic and the SOC should also be (cautiously) open to simply removing alerts that have high false positive rates.
Ensuring that all detection use-cases are subject to review will also provide a valuable opportunity to assess the capability of the SOC and identify areas for improvement.
However, it's important to note that just because a use-case has not fired, does not mean it is invalid. This is where exercises (simulating attacks) can also be incorporated into the use-case lifecycle to test the SOC capability.
Testing use-cases
As with most development practices, it is important to know if the product works. The same is true for detection use-cases.
Typically, the process of creating use-cases will involve testing, ensuring that the correct fields and formats are used to parse and display data. However genuine attacks are difficult to detect, and may not appear in the logging data as you might expect. This is why testing your use-cases is vital.
A common approach is to perform, red or purple team exercises of the estate you monitor:
Red team exercises are where an authorised individual or team simulate an (non-destructive) attack on your estate. The output of these exercises would typically be a report of any vulnerabilities found in the target system. This type of exercise is valuable in testing your capability as ideally your SOC would have detected all of the red team activity. If not, it would certainly highlight the areas in which you need to improve.
Purple team exercises are where the same type of team would perform similar attacks but also work with your SOC to iteratively improve your capability as any gaps are discovered.
There are also open source tools available if you wish to implement your own testing platform, but there are risks associated with running tests on production environments, i.e. inadvertently changing the way the production systems work or worse, preventing them from working at all.
Note: Make sure you adequately protect your catalogue of detections. If these are discovered by malicious parties, they could be used to identify gaps and exploit your system.
Frameworks
Frameworks such as Mitre Att&ck can be extremely helpful in ensuring that you have coverage for a range of attacks. When combined with threat modelling, Mitre Att&ck can help you reason about cyber attacks and ensure that your system includes appropriate and proportionate defences.


