Guidance
Machine learning principles
These principles help developers, engineers, decision makers and risk owners make informed decisions about the design, development, deployment and operation of their machine learning (ML) systems.
Our advice & guidance covers a broad range of topics
Resources for individuals and organisations in the UK who have experienced an online scam or cyber attack.
Find a range of products & services from NCSC and certified 3rd party suppliers
Working with industry, government and academia to support the next generation of researchers, students and cyber security professionals
All the latest information to help you keep track of what's happening
Page 14 of 22
You know what expected inputs to your model look like, and have criteria to identify anomalous inputs.
You can audit use of the system and its inputs and outputs.
You have appropriate log data to allow you to investigate a compromise of your system, even if not identified immediately.
You have a well-defined set of procedures for managing security incidents.
Understanding how users are querying your model and flagging unusual behaviour for investigation can help you identify and prevent attacks. Many attacks, including model inversion, membership inference and denial of service, are conducted by querying a model repeatedly, often at a much faster rate than a typical user. Repeated querying is also used in the creation of adversarial examples to be used in model evasion attacks. With the rapid adoption of LLMs prompt injection attacks have been developed, where specific queries to online APIs are able to bypass set restrictions and allow interactions with models in unintended ways.
Understanding how suspicious behaviour might look different to expected behaviour is key. The criteria for 'suspicious' behaviour will be different for each system and you should use your understanding of your system and its intended use, to develop the criteria. This should be clear from the start of the development process and include an understanding of the expected input space. For example, there are known suspicious behaviours related to prompt injection attacks, so defensive measures can be designed accordingly.
Flagged behaviour should be retained for audit and reviewed regularly for evidence of suspicious behaviour. The NCSC has guidance about logging for security purposes, and CISA now maintains Logging Made Easy software. Note that human reviewers can be used as part of routine audits or to examine behaviour identified as suspicious. When the review contains personal or sensitive data, appropriate guidelines and security procedures must be followed for handling and viewing data.
The actions your system takes in response to detecting unusual behaviour will vary. Common responses may include limiting (throttling) the rate of queries that can be sent to your model and implementing appropriate pre-processing, to heighten the threshold for an adversarial user developing a closed box attack.


