Machine learning principles
Pages
Page 10 of 22
2.3 Manage the full life cycle of models and datasets
Goals:
-
You 'version-control' your dataset so you can (if required) roll back to a known good state.
-
You have considered what metadata should be captured, and for what purpose.
-
You know the intended use of the dataset, including where it can and can't be used.
-
You know when your data needs to be deleted.
-
You are aware of any biases in your data.
-
You know what data you hold, where it is stored and who can access it.
-
You recognise the need to manage technical debt as early as possible.
Why is this important?
Data drives the development of an ML model and impacts the model's resulting behaviour. Tampering with a dataset can therefore have a significant impact on a system's integrity. However in many workflows, datasets and models are unlikely to remain static throughout the development and wider life cycle of an ML solution. This can make it difficult to distinguish nefarious changes from legitimate updates.
It's therefore crucial that system owners implement a framework for monitoring and recording changes to an asset and its metadata throughout its life. Good documentation and monitoring strengthens your ability to respond when a dataset or model has been compromised (for example, through poisoning). It's worth noting that although datasets are used to produce models, creating an inherent connection between model and dataset metadata, they should be treated as separate entities; a dataset is not necessarily exclusive to a single model and a model can be trained on multiple datasets.
Version control allows changes to be tracked and rolled back, and for metadata to be generated for confidence checking. Tracking changes between versions enables scrutiny between releases and updates and allows developers to understand and mitigate attacks or accidents.
The most significant metadata will vary from asset to asset, and it's for the asset owner to decide what is important to record. The format of this metadata should however be standardised because information captured in a consistent way is easier to share and communicate, allowing users to be better informed and make better judgements, which ultimately improves security.
Collecting data in a machine-readable format has several advantages. One is that a digital catalogue for datasets and models can be created, allowing users to filter and search depending on requirements. Machine-readable data can also be consumed by other programs that can be used as security or monitoring solutions throughout the operational stage of the life cycle. This can help implement the tracking of model behaviour, monitoring and logging user queries as outlined in Principle 2.2.
Good dataset documentation enables more trustworthy sharing of datasets. Hashing and digital signing of a dataset can ensure that recipients of a shared dataset can verify the authenticity of its contents.
How could this principle be implemented?
Use version-control tools to track and control changes to your software, dataset and resulting model
The size of your project assets (eg, datasets) may limit your choice of technologies. There are a range of tools to choose from, ranging from the general purpose (Git) to the more ML-specific (Data Version Control (DVC), MLflow, Comet, Aim). It will be for your development team to select the best for your use case, but it should be able to:
-
track which users make changes to datasets and models with full details of any modification (including the author, time and date)
-
allow for reviews before changes to an asset are made
-
'roll back' in case of a security incident where an older version of the asset is required
The UK government has advice for maintaining version control in coding and advice specifically around using Git and GitHub.
Use a standard solution for documenting and tracking your models and data
The data from a Model Card and a Data Card is complementary. Whilst the former contains information about the model, the latter contains a summary of the data’s life cycle that the model was trained on. The ML-BOM (a variant of the software bill of materials) uses dataset and model-related metadata to form a single, comprehensive, and updatable document that is designed to support transparency and accountability.
Track dataset metadata in a format that is readable by humans and can be processed/parsed by a computer
Technical implementations for tracking datasets include a data catalogue or dictionary, or implementation as part of a bigger solution (for example, as a part of a 'data warehouse' or a version-controlled database). The metadata you track on your dataset should be decided based on your specific application. Metrics that can benefit security may include:
-
a description of how data was collected
-
the sensitivity level of data
-
key data metrics (eg, RGB values over a set of images)
-
the dataset creator or maintainer
-
intended use of the data and any known limitations
-
retention time of the data
-
recommended retirement/destruction method of the data
-
aggregated stats about the data (eg, if labelled, count of each label)
A popular method to provide transparent and human-centred dataset documentation is the use of ‘Data Cards’, which provide structured summaries of essential facts about ML datasets across a project's life cycle.
Track model metadata in a format that is readable by humans and can be processed/parsed by a computer
The 'Model Cards' format is becoming an increasing popular framework to track versions as ML models are updated or trained with different datasets throughout development or continual learning through operation.
Regardless of the chosen tool, useful metadata to track for security purposes may include:
-
the dataset on which the model was trained
-
the model creator/point of contact
-
intended scope of the model and its limitations
-
secure hashes of (or digitally signed) trained models
-
retention time of the dataset (as per the dataset metadata)
-
recommended retirement/destruction method of the model
Model cards stored alongside the model in a searchable (indexed) format allows developers to find an appropriate model easily and to compare models. GCHQ have released a framework ('Bailo') on GitHub to allow this.
Ensure each dataset and model has an owner
Development teams should include roles responsible for overseeing and owning risk. Depending on the size of the team, it may be useful to establish a specific role with responsibility for managing digital assets, as part of an MLOps approach.
Whether or not a specific role is appropriate, every artifact created should have an owner responsible for ensuring its life cycle is managed securely. Ideally, this should be someone involved in creating the asset. Their contact details may be captured in the asset's metadata, although in doing this, consider the security implications of sharing personal information, especially if in the public domain.
Following these best practices can help minimise technical debt.


