Machine learning principles
Pages
Page 17 of 22
4.2 Appropriately sanitise inputs to your model in use
Goals:
-
You handle and process inputs to your model in ways that help it operate securely.
-
You protect sensitive data in line with appropriate legal, regulatory and ethical guidance.
Why is this principle important?
Appropriate sanitisation or pre-processing can help protect your model from attacks that rely on specially crafted inputs to create an adversarial effect. An example of this would be an evasion attack that used tiny changes to an input image to fool an object detection model. In some cases, sanitising the image using a simple jpeg compression could prevent this type of attack from working successfully.
If your model is trained using sensitive data (either prior to release or by CL once in use), then an attacker could try and extract this information. Properly sanitising or anonymising data could reduce the impact of a data extraction attack, and potentially make an attacker less likely to conduct one in the first place.
CL can cause particular complications around protecting privacy, as it requires collecting data directly from users.
Implement tracking and filtering of data
As a part of your data collection and MLOps processes, filters can be used for sanitising and cleaning incoming data. This can prevent your model being exposed to data that is unwanted or deliberately malicious. Consider introducing a standard set of transformations and augmentations that will standardise an input for model retraining when in use.
The filters that you require will depend on the application type, where you expect data to come from, and how much you can trust that source (in the case of CL, are they an authorised, vetted user?).
Examples could include:
-
using hateful word detection in an NLP or LLM-based application
-
taking pixel value averages for images
-
identifying outlying data points
What happens to filtered data will depend on the application and it will likely need regular review by humans. Filters will also need regular review and may need updates if adversaries find ways to bypass them.
Implement out of distribution detection on model inputs
Your development team should choose the best OoD detection for your application. Options include using maximum softmax probability and temperature scaling. If you have a complex model with large feature spaces that are OoD, it may be worth exploring whether you can factor the OoD distance into the model’s confidence scores. In some cases, this could act as detection mechanism for adversarial attacks. Further information on OoD detection can be found in this article from Encord.
Use appropriate techniques to anonymise user data
This allows you to collect data while protecting privacy, reducing security concerns. Data anonymisation techniques include:
-
Differential privacy: using statistical techniques to ensure privacy without affecting the dataset's overall statistical integrity.
-
Data masking: a range of techniques that represent a point in a different way that’s usually unreadable to a human, such as encryption, hashing or data obfuscation.
-
Generalisation/aggregation: 'zooming out' of the data to anonymise individual contributions while keeping overall trends.
-
Data swapping: swapping attributes/data points between entries while maintaining the underlying statistics of the dataset.
-
Pseudonymisation: removing (and keeping separate) attributes of the data such that a data point can no longer be attributed to a specific entry without the removed information.


