Skip to main content
Guidance

Machine learning principles

These principles help developers, engineers, decision makers and risk owners make informed decisions about the design, development, deployment and operation of their machine learning (ML) systems.

Page 10 of 22

2.3 Manage the full life cycle of models and datasets




The UK government has advice for maintaining version control in coding and advice specifically around using Git and GitHub.

Use a standard solution for documenting and tracking your models and data

The data from a Model Card and a Data Card is complementary. Whilst the former contains information about the model, the latter contains a summary of the data’s life cycle that the model was trained on. The ML-BOM (a variant of the software bill of materials) uses dataset and model-related metadata to form a single, comprehensive, and updatable document that is designed to support transparency and accountability.

Track dataset metadata in a format that is readable by humans and can be processed/parsed by a computer

Technical implementations for tracking datasets include a data catalogue or dictionary, or implementation as part of a bigger solution (for example, as a part of a 'data warehouse' or a version-controlled database). The metadata you track on your dataset should be decided based on your specific application. Metrics that can benefit security may include:

  • a description of how data was collected

  • the sensitivity level of data

  • key data metrics (eg, RGB values over a set of images)

  • the dataset creator or maintainer

  • intended use of the data and any known limitations

  • retention time of the data

  • recommended retirement/destruction method of the data

  • aggregated stats about the data (eg, if labelled, count of each label)

A popular method to provide transparent and human-centred dataset documentation is the use of ‘Data Cards’, which provide structured summaries of essential facts about ML datasets across a project's life cycle.

Track model metadata in a format that is readable by humans and can be processed/parsed by a computer

The 'Model Cards' format is becoming an increasing popular framework to track versions as ML models are updated or trained with different datasets throughout development or continual learning through operation.

Regardless of the chosen tool, useful metadata to track for security purposes may include:

  • the dataset on which the model was trained

  • the model creator/point of contact

  • intended scope of the model and its limitations

  • secure hashes of (or digitally signed) trained models

  • retention time of the dataset (as per the dataset metadata)

  • recommended retirement/destruction method of the model

Model cards stored alongside the model in a searchable (indexed) format allows developers to find an appropriate model easily and to compare models. GCHQ have released a framework ('Bailo') on GitHub to allow this.

Ensure each dataset and model has an owner

Development teams should include roles responsible for overseeing and owning risk. Depending on the size of the team, it may be useful to establish a specific role with responsibility for managing digital assets, as part of an MLOps approach. 

Whether or not a specific role is appropriate, every artifact created should have an owner responsible for ensuring its life cycle is managed securely. Ideally, this should be someone involved in creating the asset. Their contact details may be captured in the asset's metadata, although in doing this, consider the security implications of sharing personal information, especially if in the public domain.

Following these best practices can help minimise technical debt.

Published

Reviewed

Version

2.0