Machine learning principles
Pages
Page 11 of 22
2.4 Choose a model that maximises security and performance
Goals:
-
You understand that different models are suited to different tasks and the importance of benchmarking a model’s suitability for a given task.
-
You understand that different models have different vulnerabilities, and the risks associated with using MLaaS and foundation models.
-
The complexity of your model is justified for meeting your performance requirements.
-
You understand the potential trade-off between interpretability and predictive power and the effects this can have on security.
-
You understand the potential security issues that overfitting and underfitting models can cause.
Why is this important?
Once you are confident that the task at hand is most appropriately addressed using AI, it is important to consider which model should be used. You should consider security as well as performance; choosing a sub-optimal model may result in poor performance and be open to exploitation by attackers.
It’s important to properly evaluate whether to use externally pre-trained models or machine learning as a service (MLaaS) platforms. While their use eliminates many of the technical and economic costs of implementing large resource heavy models, they come with their own security characteristics which may leave you vulnerable to transfer attacks that have been designed for the base model.
Model capacity describes the complexity that a model can express and is dictated by the architecture choice (such as a neural network versus a decision tree) and size of the model (ie the number of tuneable parameters). For a given model to perform well, it's important that its capacity fits the dataset size and variety. A model with capacity that is too low for a task won't generalise well and will perform poorly. In contrast, a model with a capacity that is too high for the dataset may perform and generalise well, but may be more vulnerable to exploitation by (for example):
-
increasing the likelihood of training data memorisation
-
giving attackers extra capacity in which to hide malware
-
increasing the proportion of feature space in which the model behaves unreliably (if the training data does not adequately cover the feature space)
To validate your model for security or robustness, it's important to understand (and be able to explain) its behaviour. This allows you to:
-
detect anomalous behaviour
-
better predict how the model will react to Out of Distribution (OoD) inputs (that is, inputs that don't appear in its training data)
-
understand when you may require downstream rules or controls to constrain the output or effects of your model
Some model architectures are harder than others to interpret and understand, so your design choice must be justified (see the illustration below for details). The use of complex architectures such as neural networks should be a deliberate choice based on the requirements of your project.

The relationship between interpretability and predictive power of various machine learning architectures.
How could this principle be implemented?
Consider a range of model types on your data
Assess the performance of a range of architecture types (where performance is likely to include robustness), using a bottom-up approach to selecting your model architecture. Start with classical ML and highly interpretable techniques where possible, rather than diving into the latest deep learning models. Model selection is the data scientists/ML practitioner's job, and they must be focused on fulfilling requirements with no bias towards the latest algorithms.
Consider supply chain security when choosing whether to develop in house or use external components
For example, when assessing the suitability of MLaaS platforms or pre-trained models from model hub pages (such as Hugging Face), ensure that pre-trained models come from reputable sources. Use vulnerability scanning tools (such as Pickle scanning) to check for potential threats.
Review the size and quality of your dataset against your requirements
Rules of thumb for a evaluating the required size for a dataset will depend on your specific context and may include
- for regression analysis: the 1-in-10 rule, 10 cases per prediction
- for computer vision: 1,000 images per class
- for classification: the learning curve equation, as explained in this Dataquest tutorial
Common metrics for evaluating a dataset's quality include completeness, timeliness, validity, consistency, integrity and balance between classes. Where you have insufficient data to train a model from scratch, possible approaches include gathering more data, using transfer learning, augmenting existing data and synthetic data generation. Selecting an algorithm that runs well on limited data may also prove useful. The UK’s Defence Science and Technology Laboratory have guidance on ML techniques for limited data problems.
Consider pruning and simplifying your model during the development process
Simple models may achieve sufficient performance in many applications, but if a complex neural network model is required, a range of techniques exist for reducing its size. Many fall under the umbrella of 'pruning' (as described in this PyTorch tutorial), which involves removing unnecessary neurons or weights after the model is trained. Be aware, however that pruning can impact performance and has mostly been explored in an academic research context.


