Skip to main content
Guidance

Machine learning principles

These principles help developers, engineers, decision makers and risk owners make informed decisions about the design, development, deployment and operation of their machine learning (ML) systems.

Page 8 of 22

2.1 Secure your supply chain




It's important to understand the limitations and the risks posed, as discussed in the explainer on synthetic data produced by the Alan Turing Institute and the Royal Society. The Defence Science and Technology Library (DSTL) publish a handbook ('Machine learning with limited data') which recommends approaches for small amounts of data, and for large amounts of unlabelled data.

Reduce the risk of insider attacks on datasets from intentional mislabelling

Ensure the level of vetting for your labellers is appropriate for the severity of impact that mislabelling could have. Refer to any industry specific guidance on personnel security and vetting, and in situations where it's not feasible to vet labellers, use processes that can limit a single labeller's influence (such as breaking a dataset into segments, ensuring that a single labeller never has access to an entire set). 

Data labelling will be unique to your application, so it's important that you provide guidance and training to your labellers. Google offers useful advice on doing this. Consider whether you may benefit from the use of labelling software or a labelling service.

Published

Reviewed

Version

2.0