Machine learning principles
Pages
Page 5 of 22
1.3 Minimise an adversary's knowledge
Goals:
-
You understand the risks of disclosing information, and how releasing unnecessarily detailed information could help an attacker.
-
You understand that information on code libraries, pre-trained models or open source datasets may also be an asset to adversaries.
-
You can make a balanced assessment of the benefits and risks of sharing information about your systems.
Why is this important?
Sharing information about system designs, development processes and research findings can be beneficial to the whole ML user community. Similarly, responsible disclosure following an attack is good practice and is likely to improve ML security throughout the industry. However, reconnaissance is often the first stage in an attack, and publishing too much information about model performance, architectures and training data can help an adversary to develop attacks.
For example, If you make it known that you are developing a model by fine-tuning an open source base model (transfer learning), you may make yourself more vulnerable to an attack which has been developed on a surrogate model with the same parent model as the victim. A similar risk arises from the use of foundational models (for example large language models); if it is known that a variant of a certain open source LLM is being used to drive a particular service, then the knowledge about the parent model can be used to devise a transfer attack.
When deciding whether to release information, consider the balance between the motivation for sharing (marketing, publication or improving security practices) while protecting core system, development and model details. This balance requires understanding the implications on your system vulnerability of releasing information.
As a part of this balance, you should also consider where a model or system may be used in the future, rather than just in its current application or format. Many elements of the ML industry have roots in research and academia, where there is a culture of knowledge sharing. This means that details of many models in operation may already be public knowledge.
How could this principle be implemented?
Develop a process to review information for public release
Creating the right process starts with identifying what type of information needs to be considered. Examples may include marketing material, privacy policies and an organisation's contributions to academic literature or open source software. Assess the risk that publication could pose to your system. Seek views from a range of backgrounds (for example ML practitioners, security, software developers, non-technical subject-matter experts) to build up a comprehensive picture of the potential impact.
Your process will reflect your application and risk appetite, but it's likely to include an assessment of:
-
the information to be released against known vulnerabilities such as those in MITRE ATLAS attacks that were carried out using publicly available information (as discussed in the ProofPoint Evasion and GPT-2 Model Replication case studies)
-
the information to be released against the vulnerabilities in your system
-
how releasing the information will benefit your organisation, system or the wider community
-
whether the intended message of publication is worth the risk of any potential negative security consequences
Brief the risks to non-technical staff
Make sure that all staff (not just staff in technical roles) involved in releasing information to the public are trained on the potential impact of releasing the material. Make them aware of the required processes for releasing different types of material mentioned above. When briefing staff, the aim should be to make sure they understand what material should be reviewed, and how to do this.


