Machine learning principles
Pages
Page 8 of 22
2.1 Secure your supply chain
Goals:
-
You understand your supply chain and have sufficient transparency, traceability, validation and verification processes in place to ensure it can be trusted.
-
You understand that your data and its labels determine the logic of your model.
-
You understand that the risk of supply chain contamination heightens when you use models that don't have a transparent creation method.
-
You understand the risk posed by insider threats.
Why is this important?
An ML model is only as good as the data it is trained on. Whoever creates, processes and labels your data therefore has direct influence on your model's behaviour. Data collection, cleaning and labelling is an expensive process and in practice, people often rely on third party datasets. This has benefits, but it also introduces a threat vector. Data could be benignly but badly labelled, or an adversary could deliberately mislabel instances or insert triggers (a technique known as data poisoning).
When a model is trained on a poisoned dataset, its performance can degrade with serious consequences (as the poisoning attack on Microsoft's Tay chatbot illustrates). Other poisoning attacks can introduce targeted backdoors which could harm the integrity or availability of a model's outputs.
Supply chain security is even more important when you are using other people’s models. In one example, researchers demonstrated how an LLM model could be made to produce erroneous output, then uploaded to a public model hub, disguised as a clean implementation of another popular model. Model files can also be exploited without affecting the model itself, for example by exploiting the model serialisation process or even the model compiler.
How could this principle be implemented?
Verify any third-party inputs are from sources you trust
You need to consider the security of your supply chain when acquiring assets, including models, data, labels and software components. Require suppliers to adhere to the same standards your organisation applies to other software. Follow the NCSC's supply chain security guidance, which advises understanding your suppliers and their security posture, and ensuring they are aware of your security expectations. You can also follow advice provided by the Supply-chain Levels for Software Artifacts (SLSA) framework.
Understanding your external dependencies can help you identify vulnerabilities and risks. For example, by generating a Machine Learning Bill of Materials, you can assess all the open source and third-party components present in your ML solution.
Use untrusted data only as a last resort
If you must use untrusted data, you should be aware that techniques to find poisoned instances in a dataset (or to prevent them from having an impact) are mainly tested in academia. We therefore suggest their use only as a last resort when requirements prevent you gathering more trusted data. You should carefully evaluate their applicability and impact on your application and development process. There are also scanning tools that check code and data integrity, though these to remain imperfect. Hardening techniques can also be applied to models to detect and/or mitigate the effect of backdoors introduced by data poisoning.
Consider using synthetic data or limited data
There are a range of techniques to help you train ML systems on limited data, rather than data from untrusted or less secure sources. Some of these come with their own challenges, such as:
-
generative adversarial networks (GANs) amplifying and recreating statistical distributions in the original data
-
game engines introducing specific characteristics
-
data augmentation being limited to manipulating only the original dataset
It's important to understand the limitations and the risks posed, as discussed in the explainer on synthetic data produced by the Alan Turing Institute and the Royal Society. The Defence Science and Technology Library (DSTL) publish a handbook ('Machine learning with limited data') which recommends approaches for small amounts of data, and for large amounts of unlabelled data.
Reduce the risk of insider attacks on datasets from intentional mislabelling
Ensure the level of vetting for your labellers is appropriate for the severity of impact that mislabelling could have. Refer to any industry specific guidance on personnel security and vetting, and in situations where it's not feasible to vet labellers, use processes that can limit a single labeller's influence (such as breaking a dataset into segments, ensuring that a single labeller never has access to an entire set).
Data labelling will be unique to your application, so it's important that you provide guidance and training to your labellers. Google offers useful advice on doing this. Consider whether you may benefit from the use of labelling software or a labelling service.


