Machine learning principles
Pages
Page 13 of 22
3.1 Protect information that could be used to attack your model
Goals:
-
You understand what information is available to users, and how attackers could exploit this.
-
You understand the trade-off between transparency and security to make informed judgements about the potential consequences of an attack.
-
Your model provides users with useful answers but doesn't reveal an unnecessary level of detail.
-
You create system configuration options that provide defences against common threats.
Why is this important?
Knowledge of your model can enable prospective attackers to create better performing attacks against it. The range of this knowledge can be described on a scale between 'open box' (when an attacker has complete information about a model's architecture, weights and biases) through to 'closed box' (when an attacker has no prior knowledge, except for the ability to query the model and view its decision). This knowledge may be derived directly via access to the model, or inferred via knowledge of its data input pipeline. For example, reconnaissance of pre-processing steps, such as the shape of input data, gives an attacker information that can be used to help move closer to an open box attack.
An example of an information extraction attack is model stealing, which involves the creation of a substitute model to imitate a target, such as an ML as a Service (MLaaS) model. Data is collected by querying the target then inferring information about the model's architecture and training parameters.
During operation, providing users with visibility of a model's output can allow for quick diagnosis that a system is behaving unexpectedly. The creation of many attacks on ML models relies however on the same information; a model's output for a given input. If an attacker receives a more detailed output, it can make a successful attack more likely. It's therefore desirable to find a balance between protecting key information, while still allowing for 'sanity checking'.
A suitable balance between transparency and security will depend on the specific system application. It's likely that your system will have different classes of user, such as a ‘standard’ user querying the system, an engineer servicing the system, and a system admin diagnosing errors. Each class requires a different level of detail about the model's behaviour and model outputs should be tailored appropriately. This can be seen as a type of role-based access control.
How could this principle be implemented?
Ensure your model is appropriately protected for your security requirements
Unless further detail is needed by another process or the end user, the model should be limited to providing the required output or prediction. This limits the potential of adversarial attacks by obscuring accuracy scores and weights. You should be aware that certain adversarial techniques allow forms of model inference or evasion based only on the model output (although this is technically more challenging for an adversary.)
Due to the recent popularity of LLMs, prompt injection attacks have become a concern. There are a number of mitigations that can be put into place including limiting what can be entered as a prompt. However, it is recognised that such mitigations cannot eliminate the likelihood of an attack, so you should exercise caution when using LLMs. You may still want to consider access controls here (such as role-based access control), enabling different levels of insight depending on user authorisation.
Use access controls across different levels of detail from model outputs
Consider what level of detail each user requires. Can you represent data in a way that allows for a user to quickly check the system's behaviour without giving easy access to detailed output information? Two popular implementations of access controls are role-based access controls (RBAC) and attribute-based access control (ABAC).
Limit higher-fidelity access to model outputs to authorised users in particular roles, such as trusted maintainers and developers. For user roles that require transparency (but aren't necessarily fully trusted) consider displaying information in summary or graphical format rather than allowing raw values to be accessed. If access to the values is required, can they be with lower accuracy, or provided at a slower rate (ie lower frequency of queries)?
Configuration settings (such as model output settings per role) should be assessed in conjunction with the benefits they derive, and any security risks they introduce. Ideally, the most secure configuration will be integrated into the system as the only option. When several configuration options are necessary (ie a configuration per role), the default option should be broadly secure against common threats (that is, secure by default).


