Public content provenance for organisations
Pages
Page 4 of 6
3. Provenance: selecting suitable systems and technologies
The subject of content provenance isn't entirely new but advances in technologies, such as Generative AI, are driving requirements for it to evolve even faster.
Frameworks which offer ways of structuring provenance systems are still being established.
There are multiple facets to the provenance challenge which will require different approaches. One approach may not necessarily solve all of an organisation’s content provenance requirements. Some examples of the current provenance challenge include synthetic media labelling, provenance of digital source media, deep fake detection and provenance of aggregated content.
Organisations will have to identify a framework relevant to their needs. The key aspects to consider when selecting a framework include:
- How trust in the provenance record is established – does it use cryptographic methods such as trusted timestamps (see 3.2.1) and cryptographic identities (see 3.2.2) to secure integrity?
- The requirement to identify who will perform the third-party notary function – they have to be trusted by both the requester and the verifier.
- How verifiers, for example the courts, journalists, or a member of the public, can verify provenance. Are the mechanisms simple and understandable?
Organisations will have their own content provenance requirements but should be mindful of the rapidly evolving requirements and standards in public provenance infrastructure. They should consider standards used in their specific solution to ensure provenance functionality, such as verification work at scale.
In addition to choosing a provenance solution which meets its specific objectives, an organisation will need to decide which technologies to use.
3.1 What to consider when selecting provenance systems
Provenance systems vary in complexity, cost and effectiveness and organisations will choose their solution to meet their specific objectives. It is also important to consider that digital provenance technologies are in their infancy and that organisational requirements will inevitably evolve. For this reason, an organisation may choose to implement partial or iterative solutions.
The following section provides information on the aspects to consider when choosing provenance methods.
3.1.1 Source of trust
What is the source of trust for the content provenance record? Organisations may use internal services but will need to consider ways to mitigate the perception of 'self-signing' the provenance record. This challenge can potentially be addressed by using third-party attestation services. Organisations will need to consider the reputation and stability of third-party organisations used for establishing the provenance record.
3.1.2 Extent of provenance record
How far back does the provenance record go? At a minimum, it should trace the content back to its publication date, and identify whether the information came from a real-world device or was generated by an AI system. Ideally, the provenance should be traceable all the way back to the creation of the original source material and include provenance information about other components it contains, such as images.
3.1.3 Ease of verification
How simple is it to verify the provenance of a content item? In most cases the verifier will be a member of the public. The verification mechanism must be simple to use and yield an easily understandable and accurate provenance record.
3.1.4 Cost of providing provenance
How much does it cost to provide the provenance record? The organisation must be able to sustain the costs.
3.1.5 Strength of provenance claim
How strong are the provenance claims? Can facts about the identity and time claims stand up to scrutiny? Cryptographic validation by other parties can strengthen the claims and improve public trust in the content’s provenance record.
3.1.6 Duration of the provenance claim
How long will the provenance record need to exist? If it's in the range of years or decades then consider the sustainability of both the content store and the verification mechanisms.
3.1.7 Utility of the provenance
How does the provenance mechanism aid in reducing errors or distortion of an organisation's information? Does the mechanism aid the public in making decisions about the organisation’s content? Other information correction measures may be more effective for an organisation’s specific challenges.
3.1.8 Redress requirements
How is inaccurate information corrected? All countries have established legal mechanisms for responding to at least some inaccurate information claims against organisations in the form of libel laws. Most countries have laws in place to address copyright and trademark infringement issues. These and other laws can be used by organisations to seek redress for inaccurate information about them.
In some cases, such as copyright, there are very structured requirements for identifying infringing material and notifying hosting services to remove it, such as labelling and deploying automated processes for submission and response. Existing and potential future legal remedies and processes should be considered as well as the cost and time required to use the redress mechanisms.
3.1.9 Privacy considerations
Can privacy of individuals be addressed? Identity of actors is an important provenance detail but it is not always possible to use it such as where there may be risk to life, reputation or other concerns of individuals providing content. In some cases it may be required by law to shield an individual’s identity.
3.2 What to consider when selecting content provenance technologies
In addition to choosing a provenance solution which meets its specific objectives, an organisation will need to decide which technologies to use.
Technologies that may be relevant for an organisation include:
- cryptographic integrity mechanisms, such as public key infrastructure (PKI)¹ identities, hashing, and trusted timestamps, which can be used to bind together parts of the provenance solution to ensure the veracity and integrity of provenance records
- authentication for devices/software, individuals, and trust anchors², which is an essential part of establishing accountability in the provenance record
- decentralised storage, which can help:
- address the continuity challenges with content and records when organisations are eventually disbanded
- ensure that one party does not have full control over the digital content or ledger records
- tamper-proof ledgers, which address the challenge of permanence in the provenance record by creating records that are impossible to alter without a record of the alteration, and are independent of the content
Consideration should also be given to which parties implement the various technologies, to maximise the trust created. Organisations that create or 'self-sign' their own provenance record are unlikely to see improvements in the trust of their content.
¹The architecture, organisation, techniques, practices, and procedures that collectively support the implementation and operation of a certificate-based public key cryptographic system.
²An authoritative entity for which trust is assumed. In the case of provenance technology this definition can be more narrowly defined as an entity that a party checking provenance can rely on for other-party verification of digital content. (The trust anchor's function is similar to that of a notary.)
3.2.1 Trusted timestamps
Trusted timestamps are a useful provenance mechanism in that they establish a trusted timestamp for content state. When implemented properly, no one should be able to change a timestamp once it has been recorded. This concept is standardised in the RFC 3161 and American National Standards Institute Accredited Standards Committee X9.95 standard (ANSI ASC X9.95).
The mechanisms use cryptographic methods to calculate a hash of the document and the timestamp. A third-party organisation generally performs the timestamping to improve trust in the mechanism. Commercial services are available to perform this function.
3.2.2 Cryptographic identity
Cryptographic identities are part of PKI. They are bound to a private cryptographic key known only to that entity. The identity can be an individual, an organisation, a machine entity such as a device or service, or can be anonymous.
Cryptographic identities are commonly anchored in public certificate authorities. They can play an important part in content provenance since they can bind individuals and devices to content and assertions on content. This can strengthen the provenance of the content.
3.2.3 Digital ledgers (Blockchain)
Blockchain is a decentralised digital ledger technology that records transactions in a secure, tamper-proof manner. Each transaction, or block, is cryptographically linked to the previous one, forming a continuous chain. This chain of blocks provides a complete and transparent history of all transactions, making it virtually impossible to alter or manipulate without detection.
Blockchains are often implemented in a decentralised file system, meaning that they are not owned by any one individual or organisation and they have no single point of failure. Organisations can use public blockchains or they may choose to use a more private implementation, depending on specific provenance needs.
The NCSC has published guidance on the use of distributed ledger technology to aid in determining whether distributed ledger is an appropriate technology for a given scenario.
3.2.4 Web archiving
Web archiving refers to the process of collecting and preserving digital content from the World Wide Web so that it will be accessible in the future, even if the content is removed from a website. The primary goal of web archiving is to create a permanent record of web content, capturing website evolutions and online information changes. This process is invaluable for the preservation of digital media provenance because it captures digital assets' original form, context, and ownership, as well as subsequent versions. The Internet Archive Wayback Machine is an example of a general web archiving service.
The web archiving approach can be expanded into a more robust provenance mechanism using cryptographic signatures and timestamps. The archived data can be used to verify the authenticity and integrity of digital content and establish its historical context.
3.2.5 Digital watermarking
Digital watermarking is not a provenance mechanism but is included here because it is often considered for addressing digital trust challenges. Digital watermarking can be overt or covert.
- Overt watermarking entails adding a visible or easily detectable watermark to content such as images or video. It is often a pattern which the viewer can see. Editing the watermark will result in distortions to the image or video that may be detectable by the end viewer if unsophisticated editing changes are made.
- Covert watermarking entails adding a watermark the viewer cannot detect to the content. It will become distorted if the image or video is edited. Distortions will not be readily detectable by viewers but will be detectable by those implementing the watermarks.
Overt and covert watermarks may provide a means of detecting some attempts at altering digital content. Many forms of overt watermarks can be removed using modern editing software. Covert watermarks are limited in effectiveness by the small number of parties that can detect changes. These considerations may therefore limit the usefulness of watermarks in addressing digital trust requirements. However, watermarking can still add value as part of a layered defence implementation.
3.2.6 The Coalition for Content Provenance and Authenticity
The Coalition for Content Provenance and Authenticity (C2PA) is an industry organisation that aims to address the prevalence of misleading online information through technical standards. It has established an open specification for documenting and certifying the source and history of media content.
The Content Authenticity Initiative (CAI), which includes major technology and media companies, is responsible for promoting the C2PA standard. C2PA is a relatively new but major standard in the provenance space, and it is still under development.
C2PA leverages cryptographic methods to establish provenance on media content. This is organised around a manifest that is stored as part of the content. The manifest can capture information about changes to an item, including the author/editor, timestamp and location, and cryptographically bind it to the content. There can be multiple manifests stored in a manifest store reflecting the history of changes to the content. This manifest store is also known as a Content Credential (represented by the 'CR' icon). The standard leverages trusted timestamps and watermarking.
3.3 Why private provenance systems aren’t suitable for public content
Most organisations have some sort of internal versioning and logging systems to track details of changes to content. These systems are private in the sense that the systems and supporting integrity mechanisms such as PKI certificate authorities are often internal to the organisation.
A private provenance infrastructure works well for corporate and some legal requirements but is largely unusable for public provenance requirements. This is mainly because the mechanism is wholly managed by the organisation and designed for restricted internal use only. Additionally, private provenance systems rely heavily on separation of duties as the main mechanism for integrity of records.
Private provenance systems lack the visibility, transparency and accountability features necessary to make their provenance capability useful for establishing public trust in an organisation’s information. To address public requirements, organisations need to reconsider provenance mechanisms for at least some of their content.