Skip to main content
Guidance

Public content provenance for organisations

Explaining why content provenance matters and how organisations can use it to verify and protect their online information.

Page 4 of 6

3. Provenance: selecting suitable systems and technologies

The subject of content provenance isn't entirely new but advances in technologies, such as Generative AI, are driving requirements for it to evolve even faster. 

Frameworks which offer ways of structuring provenance systems are still being established. 

There are multiple facets to the provenance challenge which will require different approaches. One approach may not necessarily solve all of an organisation’s content provenance requirements. Some examples of the current provenance challenge include synthetic media labelling, provenance of digital source media, deep fake detection and provenance of aggregated content.

Organisations will have to identify a framework relevant to their needs. The key aspects to consider when selecting a framework include: 

  • How trust in the provenance record is established – does it use cryptographic methods such as trusted timestamps (see 3.2.1) and cryptographic identities (see 3.2.2) to secure integrity?
  • The requirement to identify who will perform the third-party notary function – they have to be trusted by both the requester and the verifier.
  • How verifiers, for example the courts, journalists, or a member of the public, can verify provenance. Are the mechanisms simple and understandable?

Organisations will have their own content provenance requirements but should be mindful of the rapidly evolving requirements and standards in public provenance infrastructure. They should consider standards used in their specific solution to ensure provenance functionality, such as verification work at scale. 

In addition to choosing a provenance solution which meets its specific objectives, an organisation will need to decide which technologies to use.


3.1.1 Source of trust

What is the source of trust for the content provenance record? Organisations may use internal services but will need to consider ways to mitigate the perception of 'self-signing' the provenance record. This challenge can potentially be addressed by using third-party attestation services. Organisations will need to consider the reputation and stability of third-party organisations used for establishing the provenance record.

3.1.2 Extent of provenance record

How far back does the provenance record go? At a minimum, it should trace the content back to its publication date, and identify whether the information came from a real-world device or was generated by an AI system. Ideally, the provenance should be traceable all the way back to the creation of the original source material and include provenance information about other components it contains, such as images. 

3.1.3 Ease of verification

How simple is it to verify the provenance of a content item? In most cases the verifier will be a member of the public. The verification mechanism must be simple to use and yield an easily understandable and accurate provenance record.

3.1.4 Cost of providing provenance

How much does it cost to provide the provenance record? The organisation must be able to sustain the costs.

3.1.5 Strength of provenance claim

How strong are the provenance claims? Can facts about the identity and time claims stand up to scrutiny? Cryptographic validation by other parties can strengthen the claims and improve public trust in the content’s provenance record.

3.1.6 Duration of the provenance claim

How long will the provenance record need to exist? If it's in the range of years or decades then consider the sustainability of both the content store and the verification mechanisms.

3.1.7 Utility of the provenance

How does the provenance mechanism aid in reducing errors or distortion of an organisation's information? Does the mechanism aid the public in making decisions about the organisation’s content? Other information correction measures may be more effective for an organisation’s specific challenges. 

3.1.8 Redress requirements

How is inaccurate information corrected? All countries have established legal mechanisms for responding to at least some inaccurate information claims against organisations in the form of libel laws. Most countries have laws in place to address copyright and trademark infringement issues. These and other laws can be used by organisations to seek redress for inaccurate information about them. 

In some cases, such as copyright, there are very structured requirements for identifying infringing material and notifying hosting services to remove it, such as labelling and deploying automated processes for submission and response. Existing and potential future legal remedies and processes should be considered as well as the cost and time required to use the redress mechanisms.

3.1.9 Privacy considerations

Can privacy of individuals be addressed? Identity of actors is an important provenance detail but it is not always possible to use it such as where there may be risk to life, reputation or other concerns of individuals providing content. In some cases it may be required by law to shield an individual’s identity. 


3.2.1 Trusted timestamps

Trusted timestamps are a useful provenance mechanism in that they establish a trusted timestamp for content state. When implemented properly, no one should be able to change a timestamp once it has been recorded. This concept is standardised in the RFC 3161 and American National Standards Institute Accredited Standards Committee X9.95 standard (ANSI ASC X9.95). 

The mechanisms use cryptographic methods to calculate a hash of the document and the timestamp. A third-party organisation generally performs the timestamping to improve trust in the mechanism. Commercial services are available to perform this function.

3.2.2 Cryptographic identity

Cryptographic identities are part of PKI. They are bound to a private cryptographic key known only to that entity. The identity can be an individual, an organisation, a machine entity such as a device or service, or can be anonymous. 

Cryptographic identities are commonly anchored in public certificate authorities. They can play an important part in content provenance since they can bind individuals and devices to content and assertions on content. This can strengthen the provenance of the content.

3.2.3 Digital ledgers (Blockchain)

Blockchain is a decentralised digital ledger technology that records transactions in a secure, tamper-proof manner. Each transaction, or block, is cryptographically linked to the previous one, forming a continuous chain. This chain of blocks provides a complete and transparent history of all transactions, making it virtually impossible to alter or manipulate without detection. 

Blockchains are often implemented in a decentralised file system, meaning that they are not owned by any one individual or organisation and they have no single point of failure. Organisations can use public blockchains or they may choose to use a more private implementation, depending on specific provenance needs.

The NCSC has published guidance on the use of distributed ledger technology to aid in determining whether distributed ledger is an appropriate technology for a given scenario.

3.2.4 Web archiving

Web archiving refers to the process of collecting and preserving digital content from the World Wide Web so that it will be accessible in the future, even if the content is removed from a website. The primary goal of web archiving is to create a permanent record of web content, capturing website evolutions and online information changes. This process is invaluable for the preservation of digital media provenance because it captures digital assets' original form, context, and ownership, as well as subsequent versions. The Internet Archive Wayback Machine is an example of a general web archiving service. 

The web archiving approach can be expanded into a more robust provenance mechanism using cryptographic signatures and timestamps.  The archived data can be used to verify the authenticity and integrity of digital content and establish its historical context.

3.2.5 Digital watermarking

Digital watermarking is not a provenance mechanism but is included here because it is often considered for addressing digital trust challenges. Digital watermarking can be overt or covert.

  • Overt watermarking entails adding a visible or easily detectable watermark to content such as images or video. It is often a pattern which the viewer can see. Editing the watermark will result in distortions to the image or video that may be detectable by the end viewer if unsophisticated editing changes are made.
  • Covert watermarking entails adding a watermark the viewer cannot detect to the content. It will become distorted if the image or video is edited. Distortions will not be readily detectable by viewers but will be detectable by those implementing the watermarks.

Overt and covert watermarks may provide a means of detecting some attempts at altering digital content. Many forms of overt watermarks can be removed using modern editing software. Covert watermarks are limited in effectiveness by the small number of parties that can detect changes. These considerations may therefore limit the usefulness of watermarks in addressing digital trust requirements. However, watermarking can still add value as part of a layered defence implementation.

3.2.6 The Coalition for Content Provenance and Authenticity

The Coalition for Content Provenance and Authenticity (C2PA) is an industry organisation that aims to address the prevalence of misleading online information through technical standards. It has established an open specification for documenting and certifying the source and history of media content. 

The Content Authenticity Initiative (CAI), which includes major technology and media companies, is responsible for promoting the C2PA standard. C2PA is a relatively new but major standard in the provenance space, and it is still under development.

C2PA leverages cryptographic methods to establish provenance on media content. This is organised around a manifest that is stored as part of the content. The manifest can capture information about changes to an item, including the author/editor, timestamp and location, and cryptographically bind it to the content. There can be multiple manifests stored in a manifest store reflecting the history of changes to the content. This manifest store is also known as a Content Credential (represented by the 'CR' icon). The standard leverages trusted timestamps and watermarking.


Published

Reviewed

Version

1.0