Machine learning principles
Pages
Page 6 of 22
1.4 Analyse vulnerabilities against inherent ML threats
Goals:
-
You understand potential defences and mitigations for the ML-specific attacks, and have incorporated them into your system design.
-
You use your new attack and mitigation-related knowledge to update the initial threat model created in Principle 1.2.
-
You understand the importance of evaluating the model and all design decisions throughout the development process, and evaluate your security requirements before a model goes into production.
Why is this important?
Identifying specific vulnerabilities in your intended workflows or algorithms at the design stage helps you reduce the need to mitigate vulnerabilities in an operational system (and the work and disruption that comes with that). Some of the vulnerabilities inherent to ML workflows and algorithms will be more relevant than others in your system, depending on variables such as:
-
data source(s)
-
data sensitivity
-
deployment and development environment
-
wider system application and consequences of failure
For example, for a query-based chat assistant model, allowing unconstrained user inputs could enable a prompt injection attack, letting an attacker gain unintended access to system instructions or private data. A more secure approach would be to restrict or filter user inputs and limit the model's ability to access sensitive data.
It's good practice to review decisions throughout the development process, with a formal security review before a model or system is released into production. The use of (automated) tools can make this process easier and more effective.
How could this principle be implemented?
Implement red teaming
In cyber security, red teaming means playing the role of an adversary to try and compromise the confidentiality, integrity or availability of a system and provide feedback on its vulnerability.
Some large technology companies (such as Google) now have ‘AI Red Teams’ dedicated to probing AI/ML systems for vulnerabilities. Though not all organisations will have the resources for that, applying a ‘red teaming mindset’ (without necessarily running a full red teaming exercise) can help inform security requirements and drive design decisions.
More information on red teaming can be found in the MOD red teaming handbook. Again, the emphasis is on applying a red teaming mindset rather than a full red team operation.
The attacks described in the resources mentioned previously (Microsoft's blog on Failure Modes in Machine Learning, NIST's AML Taxonomy, the MITRE ATLAS knowledge base and the BSI's AI Security Concerns in a Nutshell report) are a good foundation for thinking about ML vulnerabilities, while the Language model Vulnerabilities and Exposures (LVEs) project aims to raise awareness about vulnerabilities in state-of-the-art LLMs. You can establish detailed security requirements by considering these vulnerabilities in combination with your risk appetite.
Continue to apply a red teaming mindset at regular intervals throughout the system development cycle and whenever key design decisions are made. It's likely that you will get the best results when your team has skills and expertise in different areas, for example data science, software assurance and knowledge of use cases for your product.
Consider automated testing
Consider using automated tools to quantify and automatically test your model's security performance against known vulnerabilities and attack techniques. Running automated tests throughout the model's development cycle will help practitioners meet the required security goals. Commercial products to assess model robustness are beginning to come to market. There are also a number of open source tools that do this, which include:
-
The DARPA GARD (Guaranteeing AI Robustness to Deception) Program: a selection of tools that ‘seek to establish theoretical ML system foundations to identify system vulnerabilities, characterise properties that will enhance system robustness, and encourage the creation of effective defenses’.
-
Azure Counterfit: a generic automation layer for assessing the security of machine learning systems. It brings several existing adversarial frameworks under one tool, or allows users to create their own.
-
CleverHans: a Python library to benchmark machine learning systems' vulnerability to adversarial examples.
-
AI Verify: an AI governance testing framework and software toolkit that validates the performance of AI systems against a set of internationally recognised principles through standardised tests.
Note: It's likely that testing standards will emerge as the field matures, so it's helpful to keep up with the latest advice from government’s AI Standards Hub, an initiative dedicated to the evolving field of standardisation for AI technologies.


