John Collomosse FBCS, Senior Principal Scientist at Adobe, tells Martin Cooper MBCS about C2PA, a tamper-evident metadata standard, that’s looking to embed an image’s origin story into the picture itself.
Summary:
- Online content now originates from a decentralised creator ecosystem of creators, and trust can no longer rely on organisational reputation
- Provenance labelling, which creates an unalterable history of a piece of digital content, may be the solution to identifying manipulated or deceptive content
- Provenance labelling tells audiences clearly how content has been made, edited and shared, giving audiences power and allowing good actors to prove the authenticity of their content
- The process combines the 'three pillars of provenance'; metadata, watermarking and fingerprinting, making identification durable
- Provenance labelling can carry things like AI consent preferences so that people and systems can understand how content can be used, impacting the digital creative economy
Deep fakes, misinformation, synthetic media, counterfeits — whatever you choose to call what’s not real, digital deception has reached crisis point online. And as AI tools become more powerful, prevalent and easier to use the problem only gets worse.
The solution may reside in content provenance: a tamper-evident digital ‘nutrition label’ inserted directly into media files. By fusing secure metadata with invisible watermarking and cryptographic fingerprinting, this technology survives social media stripping to map an unalterable history of how content was created.
Far beyond just fighting fake news, this decentralised infrastructure can also support creator consent, attribution and licensing. In the future, it could provide a shared foundation for online trust, whilst also helping creators gain better attribution and fairer compensation across the internet.
Why don’t you introduce yourself and tell us a little about your career?
I am Professor of AI at the University of Surrey and Senior Principal Scientist at Adobe Research, where I lead Adobe’s content authenticity research programme. At Surrey, I founded and direct DECaDE, the UKRI Centre for the Decentralised Digital Economy, where we research how emerging technologies can improve trust, data agency and value exchange in the creative industries.
My career began in computer vision and AI research, but has increasingly focused on trust in the digital supply chain: how we know where data and media came from, how they may have changed, and how rightsholders can retain agency as their work moves through online platforms and AI systems. This matters not only for the creative industries, but for data governance and integrity in AI more generally.
What is content provenance and why is it important? What problem are you looking to solve?
Generative AI is empowering creators, but also brings risks. Fake news and misinformation are major societal challenges, and training AI using creative work has raised urgent questions around consent and compensation. We believe content provenance can help address both of these risks.
Content provenance is the story of how a piece of digital content was made. A photographer might take a snap, retouch it, and pass it to an editor who crops it for social media before a publisher releases it online. Provenance is a way of recording that history so that people and systems can understand where the content came from and what has happened to it as it passes through such a digital supply chain.
You can think of it like a digital nutrition label. We care about the provenance of our food: where it came from, and how it has been processed — why not our news? Knowing whether an image came from a camera, was edited or generated by AI, or re-used from an earlier source gives us context to make informed trust decisions about that content.
Surely spotting deepfakes is a task that AI should excel at?
A deepfake is AI content that has been created with intent to deceive, but a computer can’t detect intent — it can only detect AI use. AI detection tools are valuable in some instances, but their results can be unreliable and will always be in an arms race against bad actors. Detection itself is becoming a weaker signal because most content is now touched by AI. Even basic image operations such as upscaling and filling use AI! There is also nothing inherently deceptive about AI-generated media. Many people use AI to create and tell authentic stories. Simply labelling content as ‘AI’ tells us less and less.
Conversely, human rights organisations like WITNESS have reported that most visual misinformation is not AI, or even edited, at all. It often involves unaltered content that has been misattributed to tell a false narrative.
That is why provenance is so important. Instead of trying to classify content simply as ‘real’ or ‘fake’, or ‘AI’, provenance describes how content was made, edited and shared. It gives people the context they need to make up their own minds.
How did you get interested in provenance? Was there a moment when you thought: ‘I have to solve this’?
10 years ago, I led an academic project called ARCHANGEL with The National Archives and the Open Data Institute. It used provenance to underwrite the integrity of born-digital media, such as video records from the UK Supreme Court. The project recorded provenance information on a blockchain: a shared, tamper-evident database maintained by multiple archives: the UK, Estonian, Norwegian, Australian and US National Archives (NARA) were involved in the trial. The aim was to enable members of the public to verify the provenance of digital records when they were released. We later extended the prototype into the news domain through the open-source ‘Angel’s Wing’ browser extension, which searched for provenance records associated with photographs on the web.
That work made it clear that provenance was going to be important. Digital media no longer comes only from institutional sources such as broadcasters or archives. It now comes from a decentralised ecosystem of creators. Trust can no longer rely only on the reputation of the institution that publishes something. We need technologies to help people understand the history of the content itself.
When I later joined Adobe, several of us saw provenance as a key part of the answer. That thinking led us to create the Content Authenticity Initiative (CAI) in 2019, which today has more than 6,000 members working hard to make content provenance practical at internet scale. A major part of the CAI’s work has been to found and advocate for the Coalition for Content Provenance and Authenticity (C2PA) open standard. In simple terms, C2PA records provenance information in the metadata of digital content, so that the history of a file can travel with it.
Take us under the bonnet —how does C2PA work? How is it different from, say, EXIF metadata? What is a ‘manifest’?
A manifest is the metadata structure that C2PA uses to store provenance information. It contains provenance facts, called assertions.
EXIF can record information such as camera settings or location, but it is easy to modify. C2PA is designed to be tamper-evident. Its manifest is signed and sealed using public key infrastructure, or PKI, so people and systems can check who or what is attesting to those facts and whether they have been altered.
Assertions can record who created or published content, what actions were performed, and which other assets were used in its creation. Those source assets, called ingredients, can have their own manifests too, so forming a graph describing the history of the asset. The manifests are bound together with hashes, including a hash of the content itself, so tampering becomes evident.
The idea is that C2PA allows good actors to prove the authenticity of their content. If a politician, celebrity or news organisation routinely signs their content, then fake content purporting to be from them will not carry their signature. Signing content is also a way for creators to assert authorship and gain attribution as their work moves through the digital supply chain.
Who is using C2PA — where is it making a difference right now?
Adoption is growing quickly. Many generative AI platforms write C2PA into generated content to show that AI was used, partly driven by legislation in jurisdictions such as the EU and California. Several camera manufacturers, including Leica, Sony and Canon, have adopted C2PA in models aimed at photojournalism, and recent Google Pixel phones also support it. This lets users show their photographs and footage are real.
For you
Be part of something bigger, join BCS, The Chartered Institute for IT.
Software products are also starting to read and write C2PA metadata, including Adobe’s creative tools, such as Photoshop, Lightroom and Premiere. Much early adoption has focused on still images, but there are good examples in video too, including France Télévisions signing daily news output and the BBC’s prototype at the IBC conference.
LinkedIn surfaces provenance information when C2PA is present in posts, and Google recently announced plans to bring C2PA into search and browsers. That matters because provenance becomes more useful when people can see it where they encounter content. Still, many platforms, particularly social media, strip away C2PA metadata and break the content supply chain.
How can this scale up if platforms strip metadata?
The solution is to embed an invisible watermark in the content to make provenance metadata more ‘sticky’. The watermark uniquely identifies the content, enabling systems to recover stripped metadata from the cloud.
We developed a watermark called TrustMark, designed to survive transformations applied by social platforms. We open-sourced TrustMark to promote interoperability and adoption, although this brings security challenges: someone might transfer (spoof) a watermark onto different content. That is why watermarking is combined with content fingerprinting, which helps verify that the watermarked image really matches the metadata recovered for it. One of our research breakthroughs was to combine these three pillars of provenance - metadata, watermarking and fingerprinting - together in this way to improve the durability of provenance. The three-pillar approach is now widely adopted, including in the Code of Practice for the EU AI Act.
What’s next for content provenance?
Beyond fighting fake news, knowing the provenance of content can also help creators secure attribution for their work, and compensate them when their work is reused.
C2PA began before the rise of generative AI, but it now has relevance to questions around AI training, consent and copyright. Unlike tools such as robots.txt, which are tied to a website, C2PA can carry creator preferences around AI consent inside the manifest which travels with content. It is also possible to embed open licensing standards within C2PA. Then, wherever content is found online, people and systems can understand what it can be used for, who owns the rights, and who should be compensated if it is reused.
We see provenance as digital infrastructure for the creative economy: a stack of capabilities that starts with content authenticity, then layers on rights and licensing to unlock new forms of value creation. At Surrey, DECaDE has been pioneering this approach, including through our recent Time to ACCCT report with Sheridans and the CoSTAR National Lab. If we get this right, provenance can give creators, publishers and platforms a shared infrastructure for trust, attribution and fair value exchange in the AI age.
Finally, you’re a BCS Fellow – what advice do you have for somebody thinking about applying?
BCS Fellowship is a recognised mark of professionalism, but it is also about belonging to a wider community. I joined BCS as a computer science undergraduate, later became a Chartered Engineer through BCS, and then became a Fellow. I have found that professional standing particularly valuable as my work has increasingly involved policymakers and standards organisations.
The regular computer vision meetings held at BCS with the British Machine Vision Association have also helped maintain the visibility of the UK AI research community. I would certainly encourage people to apply when they feel they have built a strong case: it recognises your professional contribution while supporting the discipline more broadly.
Take it further
Interested in this and similar topics? Explore BCS' books and courses: