Intelligent security tools
Assessing intelligent tools for cyber security
Page 2 of 6
Defining artificial intelligence
Discover what AI is, its various methods and implementations, the kind of problems it's good at solving and those it's not well equipped for.
Before we move on to a detailed consideration of the issues related to intelligent security products, it's worth taking a moment to think about Artificial Intelligence (AI) as a technology.
This section is intended to give you a broad brush introduction to AI, its various methods and implementations, the kind of problems it's good at solving and those it's not well equipped for. We'll also highlight some of the most important vocabulary.
There are many definitions of AI in use, and very little agreement. Here we describe how we will use the term.
The essentials
How AI works in simple terms
Modern AI is usually built using machine learning algorithms. These find complex patterns in data, which can be used to form rules which are useful to us.
For example, a machine learning algorithm could find similarities in pictures of cats. If you tell it which pictures are of your cat, this can form rules that allows the system to recognise your cat. When you show it a new picture, it will be able to predict whether the new image is your cat or not, based on it's learned definition.
For us, a key part of AI is that it takes those patterns and definitions and uses them to automate a decision. For example, you could just use this to automatically file your photos. Or you could build this into your cat flap, along with a camera, with the action that it only unlocks when it recognises your cat. The decision to let cats in, and therefore your 'security', has been automated.
What can AI currently do?
AI is very good at solving well bounded problems, where the solution, and the method to find it, is completely contained within the data and feedback provided. Here they regularly excel beyond human abilities, for both speed and accuracy.
A good example of this is game playing, an area where AI has had some great successes. Chess and Go being two highlights. When playing a game, the rules form the bounds of the problem. The data from which the AI learns is all the previous games it has played. The game board becomes the AI's whole world, nothing else needs to impact the decisions it makes while playing.
Another fruitful field for AI has been image recognition and generation, where the problems are to predict which group a new image belongs to, or create a new image of that type. All the information required to classify, or create an image, is contained in the image data provided.
What is AI not so good at today?
AI is less good at solving problems where you need additional context to reach an answer, even if that context would be 'common sense' for a person. It may be possible to provide some of this context to an AI in the form of additional data, a model, or feedback from a human, but expanding the bounds of the problem is often expensive and may result in poor performance if expanded too far. At this point, the decision should be passed to a person who is able to use their knowledge of the context to reach a decision.
Cyber security often involves situations where context is important. Accessing a sensitive document might be a suspicious action for one person, but normal behaviour for another. Installing updates is essential for most of your business, but a business risk where it causes compatibility issues with critical software. This is why it is important to find a tool that can work within the context of your business. And why you may need to ensure that a person has the opportunity to step in and apply their knowledge.
The technology underpinning AI is constantly evolving, and we can expect any limitations to be continually tested and redefined. Despite its rapid pace of change, you need to be aware that there will be some areas where it is too difficult (or expensive) to develop AI capable of solving the problem.
Intelligent tools for cyber security
If we want to automate cyber security tasks at scale, we need the ability to process data in quantities beyond the capability of any human being.
To do this we have a wide range of tools available, things like AV and firewalls. Until recently, however, these have worked on manually developed rules. Now, the tools are increasingly learning the rules for themselves, using intelligent algorithms to derive these rules directly from our data.
These new rules can make the tool very powerful, but they also make them less predictable.
AI Terminology
A brief primer on some of the most common AI-related language.
Rules
A standard rule in automation may take the basic form, 'if this, then that'.
In traditional automation, we would have to identify and explicitly program the 'this' and 'that'. AI looks for patterns in data to find its own 'this'.
This is powerful because we often find it very difficult to identify and describe the 'this' in a way that the computers understand.
For example, how could we describe an image of a cat to a computer with sufficient accuracy to satisfy the condition, 'if the image is of my cat...?'
In cyber security, the condition could be, 'if the website is suspicious', 'if the image is of an employee', or 'if network activity is abnormal'. These, and many other things, are impossible to define manually.
If this then that
In many intelligent tools, we still manually program the 'that', because we know what we want the possible actions to be. In the cat example, we want to unlock the cat flap. So, the AI takes care of, 'if the image is of my cat'. Then we manually add the rest of the rule, 'then unlock the cat flap'. In cyber security this could be 'if the website is suspicious, then block', 'if the image is of an employee, then give access', or 'if network activity is abnormal, then show a warning message'.
Beyond that
Some modern AIs find their own 'that' by trial and error, using an ultimate goal as the criteria for a successful rule. For example, winning a chess match. This needs a huge number of trials to learn and usually can't be done in real time. Results are achieved by modelling the environment and simulating the trial and error process.
This is not commonly used in cyber security tools, because the ultimate goal of cyber security is difficult to measure and we cannot model all the variables necessary to allow the AI to learn through trial and error. There is also a risk that the AI will find a novel solution that is detrimental. For example, keeping all the computers switched off would fulfill the goal of preventing attacks, but our common sense tells us this is not a viable option.
Model
The rules are stored in a trained model. An AI can be made up of multiple trained models, each with their own rules, which are then combined to make the decision.
Learning and Training
Learning and training are the terms we use to describe the process whereby an AI finds or updates its rules. AI does not learn like a human. People can learn a fact by simply being told a few times. An AI has to 'see' this fact in the data at a high enough frequency to detect a pattern. This is the reason why you need such high quantities of data to train an AI. It is also why it's difficult to correct a mistake.
Machine learning algorithms are the computer algorithms that perform this learning process.
Data
The data used by an AI can take many forms, such as a traditional data set, feedback from interactions with humans, or the experience of success/failure during the trial and error process.
Data quality
The quality of an AI is highly dependent on the quality of the data used during training and operation. There are certain attributes we look for in high quality data, such as:
- Completeness - If there are blanks in the data - for example fields left incomplete - the AI will only have part of the information it needs to train or make decisions. This will lower the accuracy of any output, as some of the patterns in the data can't be discovered.
- Diversity - If all of the examples an AI uses to train look the same, then the capability for the AI to learn anything outside of those few cases is limited. A diverse data set for training will look to have a wide variety of both desirable and non-desirable characteristics for the AI to learn from. A system can only be as good as the data that it's trained on, having a diverse data set will mean that a greater variety of cases can be addressed in the trained system.
- Accuracy - The way data is presented to train an AI is important. Misrepresenting the data being processed by the AI - for example through incorrect labelling - will result in a less reliable trained system. In the live system, this will result in misclassifications and unwanted behaviours from the AI.
Intelligence
AI today, and for the foreseeable future, is not intelligent in the same way as humans. They cannot reason, feel or apply common sense to a problem.
Current AI is just a new way of processing data in our computers, where data goes in, a handle is turned and information comes out. True intelligence or 'Artificial General Intelligence' is still the stuff of science fiction.
Black boxes
A black box is part of a program where we cannot see or understand what is happening. The input data goes into the black box and the output comes out. We don't understand how the output is derived from the input. The trained models produced by many machine learning algorithms are black boxes by nature. The science simply has not advanced to the point where we can really understand the process.
The models produced by some machine learning algorithms are not black boxes - it is possible to understand how their decisions are made. In addition, there is extensive research underway in academia and industry into explainable AI. In other words, looking to make black box algorithms understandable.
Impact
Most intelligent tools will be carrying out some sort of automation, but the way this affects the real world will vary.
For example, some tools will automate the process of deriving information from the raw data, providing a warning or a suggested action. How that information is used will ultimately depend on the person who receives the information.
In other cases, the tool will be the one acting, automatically, on the information. There will be no human in the loop.
Adaptation
One of the big selling points of many intelligent tools is that they can continue to learn from new data and feedback as you use them, allowing them to adapt to new scenarios. This makes them more resilient to new situations, but it also means that your tool may behave in a way that is unpredictable, making it more difficult to gain assurance that it is working correctly.
Bias
While learning, an AI will learn any biases contained within the data provided. This could be from the original training data, or feedback provided by an operator.
Biases impact the quality of data being used to train the AI. AI trained with biased data are essentially being given erroneous patterns to find, which can prevent them from finding the real patterns that we want them to find. This can reduce the accuracy of the AI.
For example, if you give feedback on only half the problem, the AI will continue to learn based on the assumption that the other half was correct. This could potentially create problems as these issues would remain undetected.
AI can also learn the prejudices of operators. If, for example, an operator provides biased feedback against certain individuals, this may get embedded into your system, leading to discriminatory (and potentially illegal) decisions.


