# What is the TDM·AI Protocol?

DRAFT FOR DISCUSSION Last updated 2024-12-18

<figure><img src="/files/IYVMDm6C27pRNZqXHf58" alt=""><figcaption></figcaption></figure>

## Abstract&#x20;

TDM·AI is an attachment mechanism that enables creators and rightsholders to persistently and verifiably attach machine-readable usage preferences – such as an opt-out from text and data mining (TDM), automated processing, or AI training – to their digital works.

TDM·AI addresses the challenge of expressing and binding AI training and usage preferences for individual digital media assets. The protocol provides a consistent way to declare permissions or restrictions regarding the use of content in AI model training by linking these declarations to content-derived identifiers and digital fingerprints.

It uses the International Standard Content Code (ISCC [ISO 24138:2024](https://www.iso.org/standard/77899.html)) for asset identification and cryptographically verifiable  [Creator Credentials](https://docs.creatorcredentials.com/), based on [W3C recommendation for cryptographically verifiable credentials](https://www.w3.org/TR/vc-data-model-2.0/). This ensures that declarations are both verifiable and attributable to legitimate rightsholders or controllers of the content.

Originally developed in the context of Article 4 of the European Directive 2019/790 (DSM Directive), TDM·AI is applicable across jurisdictions and allows rightsholders worldwide to express usage preferences in a structured and interoperable way.

## Motivation

Given the current digital AI landscape in the context of an evolving international regulatory environment, there is an urgent need for a reliable way for content creators and other rightsholders to declare their consent or reservation to automated processing (TDM) for the purpose of training models and applications of generative AI, capable of generating text, images, and other content.&#x20;

The TDM·AI protocol aims to provide creators and other rightsholders with a simple and standardised way to make a machine-readable declaration as to whether or not their content may or may not be used for these specific purposes. The key differentiator of the TDM-AI protocol is that the AI preferences, can be resolved directly from the content-derived identifier, the ISCC code, meaning that it is easily accessible to users such as AI providers using open-source identifier technology.&#x20;

By using ISCC, the protocol ensures a reliable method of identifying content that is robust to common problems such as the loss of embedded metadata, removal of watermarks or steganographic data, or other alteration or manipulation of content.&#x20;

The use of verifiable credentials adds a further layer of trust and verifiability, ensuring that the declarations are genuine and can be traced back to the original rightsholder, depending on their privacy needs and preferences.

The TDM·AI protocol is motivated by the need to:

* Provide a clear and simple way for rightsholders to declare rightsholders' preferences with regards to training models of generative AI;
* Ensure that AI providers and other stakeholders can easily read, understand and respect rightsholders' preferences by machine technology.

## Overview

<figure><img src="/files/VgYTdklv5Hb1ULHH5Azy" alt=""><figcaption><p>Overview</p></figcaption></figure>


# Recommendations for Opt-out

Liccium's TDM·AI protocol describes the method of binding AI preferences to ISCC codes. To be effective, rights reservations must be:

* **Inseparably bound to the fingerprint of the content** allowing sharing and distribution of content;
* Easily discoverable, accessible and **machine-readable** (legal requirement in the EU);
* Functional when **content already has been shared and distributed**;
* Declarations can be **updated**;
* Resilient to content manipulation or alteration;
* **Resilient to the removal of embedded metadata** and visible or invisible watermarks;
* Containing verifiable attribution through digital signatures and certificates which authenticate the declaration's source and prevent false claims;&#x20;
* **Functional for all media types and file formats**;
* Utilising verifiable timestamps;
* Based on international standards (ISO, W3C).


# Metadata Binding

## Options for Binding Rightsholders' Preferences

When it comes to managing rightsholders' preferences for TDM in a machine-readable way, in principle, three different attachment mechanisms are discussed:&#x20;

1. **Location or domain-based metadata binding**: rightsholders' preferences for web-published content are included in robots.txt-file or in HTML/HTTP metadata of the domain, e.g. Robots.txt;
2. **Asset-based metadata binding**: provenance metadata – including rightsholders' preferences – is embedded directly into the media file, e.g. C2PA.org.
3. **Registry-based metadata binding**: ISCC fingerprints and preferences is submitted to publicly accessible registries.&#x20;

<figure><img src="/files/4ebnIPkX4XwocsMKXaAd" alt=""><figcaption></figcaption></figure>

## Why Do We Need Opt-Out Registries?

* Embedded metadata is removed from the media file.
* Content is altered or manipulated, compressed or converted into a different file format – which is where methods based on cryptographic hashing fail.
* When content is already distributed, shared and part of training sets, metadata cannot be embedded.
* Content is shared on websites beyond the rightsholder's control, e.g. on social media, so a domain-based approach cannot be applied.
* A lot of content is not publicly accessible on web domains.
* Watermarks or steganographic data can be removed from media files.


# Issues of Domain-Based Opt-out

## Robots.txt

Robots.txt has become a practical way for rightsholders to express their preferences regarding the crawling of web-published content. It is simple to implement, widely adopted, and recognised as a standard mechanism across the web.

However, when it comes to expressing AI training preferences, this approach is inadequate for many stakeholders. Location- or domain-based methods do not effectively address the realities of how AI training data is collected and used, for the following reasons:

1. Most professional creators and rightsholders do not publish their content online;&#x20;
2. Copyrighted content is republished on the Internet without authorisation;
3. Content is shared on websites rightsholders do not control, such as social media platforms or licensor's websites.

Robots.txt is insufficient as a primary mechanism for protecting creators’ rights in the context of AI training. It only applies in cases where rightsholders have direct control over a domain, cannot account for works that are republished without authorisation, and fails to cover downstream licensing or reuse of creative assets. To ensure comprehensive protection for creators and rightsholders, AI governance frameworks must adopt more advanced approaches – such as asset-level or registry-based rights reservation – that operate independently of domain ownership and reflect the realities of modern content distribution.

### Most professional creators and rightsholders do not publish their content online

Many creators in industries such as music, film, book publishing, and audiobooks do not distribute and publish their content directly online. Yet, their works may already be used in AI training datasets through other means (e.g., data bundles, unauthorised uploads). Relying solely on robots.txt disregards creators whose content is included in datasets without ever being published on a controlled website.&#x20;

A more robust framework would require AI model developers or system providers to verify whether rights reservations are put in place at the asset or work level, independent of robots.txt implementation.

### Copyrighted content is republished on the Internet without authorisation.

Content frequently appears online without rightsholders' permission due to piracy, negligence, or ignorance. While banning crawling of websites on which pirated content is published may be a step forward, it fails to address cases of unauthorised republication on legal platforms. For example, copyrighted images may circulate on forums, blogs, or social media without any indication of their original source.

Robots.txt is ineffective here because those sharing the content may not be aware of the original rights reservation or may actively disregard it without any legitimate reason.

### Content is shared on websites rightsholders do not control, such as social media platforms

Robots.txt requires control over the hosting domain, making it ineffective for content distributed on third-party platforms, such as social media or licensing partners’ websites, where creators cannot implement their individual robots.txt settings.

For instance, a photographer's image licensed to a news outlet may be shared on social media, where the original robots.txt settings are neither implemented nor enforceable – neither by the original rightsholder nor the licensee. Even when creators use robots.txt on their own websites, they cannot ensure downstream compliance. Licensing terms are often dictated by the licensee further along the distribution chain, and economic constraints may prevent the licensor from enforcing robots.txt settings effectively.

<br>


# Issues of Asset-Based Optout

There are multiple standards to embed rights and metadata inside the media file itself, such as IPTC or C2PA.&#x20;

Content creators and publishers can use apps that support the C2PA method to create and embed cryptographically verifiable metadata containing information about the asset’s creation and edit actions, copyright, licences, capture device details, and software used. This manifest may include rightsholders' preferences that enable "a human actor to provide a C2PA Manifest Consumer information about whether an asset with C2PA metadata may be used as part of a data mining or AI/ML training workflow." The assertions are designed to be hashed and gathered into a verifiable claim that is digitally signed, ensuring the integrity of the claim.

However, this 'hard-binding' of the embedded assertions and metadata within the content breaks in the following situations:&#x20;

* When embedded metadata (or the certificate) is removed from the media file;
* When content is altered or manipulated even to a small extend, as the method is based on cryptographic hashing;&#x20;
* When content is converted into a different file format, compressed or screenshotted.

However, these are very common problems when content is shared online or on social media platforms that resize or compress media files or remove embedded metadata for security and business reasons.


# Advantages of Opt-Out Registries

## The TDM·AI Protocol

The TDM·AI protocol proposes a reliable method of binding rightsholders' AI training preferences to the content-derived identifier and digital fingerprint of the media file. Creators and rightsholders can generate ISCC fingerprints directly from their content and choose the rightsholder's preference for the content. ISCC codes and the selected preferences can be publicly declared in a network of open, centralised or federated, verifiable metadata registries.&#x20;

These registries persistently bind the rightsholders' preferences to the unique identifier of the media asset (ISCC Codes) – and ‘persistently’ means: the data cannot be separated or removed. Federated directories have to be publicly accessible to discover ISCC codes, resolve associated rightsholder preferences and verify the authenticity and originality of the declarations.&#x20;

> **"Looking at the current landscape of unit-based identifiers, an approach based on a content derived identifier such as the ISCC to identify opted-out works and record opt-outs via the proposed standardised vocabulary seems viable for at least some categories of works. Such a registry would soft-bind opt-out declarations based on the standardised vocabulary to ISCC codes. This would allow AI model trainers to use ISCC codes as a look-up key to check the registry for known opt-outs."** \
> **(**&#x4F;pen Future Foundation, Considerations For Opt-out Compliance Policies by AI model Developers, [https://openfuture.eu/wp-content/uploads/2024/05/240516considerations\_of\_opt-out\_\
> compliance\_policies.pdf](https://openfuture.eu/wp-content/uploads/2024/05/240516considerations_of_opt-out_compliance_policies.pdf))

Since anyone can make a public declaration, it is important to ensure proper authentication of the source of each declaration. To increase trustworthiness, it is suggested that declaration metadata will include publicly accessible verifiable credentials (VCs), which are based on the W3C standards for Verifiable Credentials, supported by advanced and qualified certificates that properly identify creators and rightsholders. These "Creator Credentials'' serve as a means for attribution and authentication of creators and rightsholders based on social or institutional authentication, thereby increasing trust in claims and attribution. Creator Credentials provide creators and rightsholders with a sovereign, portable and interoperable way of managing their digital identities.&#x20;

All declarations are digitally signed and provided with verifiable timestamps to ensure their accuracy, validity and transparency as to when exactly the opt-out declaration was published – an aspect often overlooked in the current discussion about opt-out declarations.

## Comparison of Opt-Out Protocols

<figure><img src="/files/BnuSjxwwUDWsSKLngxtJ" alt=""><figcaption></figcaption></figure>

## Compatibility&#x20;

This table tries to provide an objective overview over existing (not all) methods to opt-out. It depends on the use case which method would work for the individual creator or rightsholder.  The use of the TDM·AI protocol does not exclude the use of other methods. It is compatible with all methods to express terms for TDM for AI. <br>


# Benefits of TDM·AI

TDM·AI utilises the International Standard Content Code (ISCC), an new ISO standard for digital media content ([ISO 24138:2024)](https://www.iso.org/standard/77899.html) and verifiable Creator Credentials, based on[ W3C recommendation for cryptographically verifiable credentials](https://www.w3.org/TR/vc-data-model-2.0/), to ensure verifiable and machine-readable declarations that include proper attribution of claims.&#x20;

The ISCC is an open system for the content-derived identification of digital media content of all media types and formats (text, image, audio, video). In 2024, ISO published the ISCC as a global standard for digital content identification.

{% embed url="<https://www.iso.org/standard/77899.html>" %}

Let's get into the details of the advantages of this solution.&#x20;

## Benefits of the ISCC

* **Content-derived Identifier** – The ISCC serves as a content-derived identifier for the specific digital media asset, allowing various parties, including users and machines, to independently generate codes only by having access to the file itself without the need of a previous registration or application process.
* **Accessible and Open-Source** – ISCC codes can be generated using open source software or applications such as Liccium. It is open-sourced with a permissive licence, lightweight, dockerised, and effortlessly deployable.&#x20;
* **Versatility Across Media Types** – ISCC can be used with all media types (text, images, video, audio) and supports a wide range of file formats.
* **Resilience Amid Metadata Removal** – ISCC allows robust identification and matching of media assets even in cases where metadata was removed, a common occurrence when content is shared online or via social media.
* **Effective When Content is Shared** – ISCC remains effective even when content is taken out of its original context and embedded metadata or attribution becomes inaccessible due to content sharing on third-party domains outside the control of the rightsholders, potentially losing creator or rights information in robots.txt files or other metadata.
* **Robustness against content changes** – ISCC is a content-derived identifier with matching capabilities for near-duplicate content, which means that it retains its reliability even if content is changed, modified or manipulated up to a certain degree.

## Benefits of ISCC Declarations

Liccium gives creators and rightsholders an intuitive platform to digitally sign and protect their original works, building trust in ownership, attribution and authenticity of digital media content. Using Liccium creators can make declarations of ISCC codes, providing access to rights and metadata:

* **Inseparably Binding Rights and Metadata** – ISCC declarations enable the secure and indissoluble binding of external metadata, rights, or other claims, along with verifiable credentials, to the content, offering substantial advantages to creators and rightsholders.
* **Machine-Readable Declarations** – ISCC declarations, in conjunction with verifiable credentials, can provide relevant and sufficient information regarding rights and restrictions in a machine-readable way.
* **Verifiable, Timestamped Data** – ISCC declarations are both timestamped and cryptographically verifiable, ensuring the integrity and temporal context of the associated information.

## Benefits of Creator Credentials

The Creator Credentials project has developed a user-centric digital identity management framework that is specifically designed to serve the unique needs of the cultural and creative industries.&#x20;

{% embed url="<https://docs.creatorcredentials.com/>" %}
Documentation of the Creator Credentials project
{% endembed %}

Creator Credentials offer a multitude of essential features and benefits that not only enhance security but also facilitate seamless interoperability across various systems and applications:

1. Creator Credentials are **self-sovereign**, which means that creators and rightsholders have primary ownership and control over their credentials. They can choose where to host, when to use, and how to share them, including controlling access levels and revoking access if needed.
2. Creator Credentials are **portable**, which means that once issued the holder can transfer credentials between different systems and platforms without loss of information or functionality. They can also be used with various applications or services as required.
3. Creator Credentials are **interoperable**, which means that credentials issued by one system can be recognised and accepted by other systems, even if they are built on different technologies or operated by different organisations. Third party applications can independently verify these credentials.
4. Creator Credentials  can also be **restricted in duration**. This means they support the ability to set expiration dates for verifiable credentials, defining a time limit on how long a verifiable credential is valid. Limited validity can be useful for credentials that expire, such as professional licence or certifications.
5. Creator Credentials  are **revocable**, which means that they can be revoked or cancelled in cases of expiration or other reasons. This is important for maintaining the integrity of the credential system and ensuring that only valid and accurate information is being shared. The revocation can be initiated by the issuer of the credential or by the individual or entity who holds the credential. It can be enforced by the systems that recognise and accept the credential.
6. Creator Credentials are **cryptographically verifiable**, which is essential for the proper attribution of rights and restrictions, ensuring transparency and accountability.
7. Creator Credentials are **resilient**. Using them in the context of ISCC declarations, even in cases of content alteration, manipulation, or removal of watermarks or metadata, Creator Credentials can still be derived from the content and its associated ISCC code.
8. Creator Credentials are **publicly accessible**, allowing for verification within the context of any public content declaration or claim.
9. Creator Credentials are **immutable and trustworthy**, meaning they cannot be altered without detection or notice.
10. Creator Credentials are **timestamped** to ensure cryptographic verifiability and transparency and to provide information about who made a declaration and when it was made, both for humans and machines.

{% hint style="info" %}
In 2023, the European Union funded a project called 'Creator Credentials' as one selected project of the NGI's Trustchain consortium under grant agreement Nr. 101093274.&#x20;

The Creator Credentials project will develop a user-centric digital identity management framework specifically designed for the cultural and creative communities. This includes a software application that can be used by media organisations to issue verifiable credentials to creators and other rightsholders. Creator Credentials will increase the trustworthiness of declarations and claims to digital media content online.

<https://trustchain.ngi.eu/creatorcredentials>
{% endhint %}


# Opt-out, opt-in and content licensing

TDM·AI defines a vocabulary and technical specification for public, machine-readable AI preferences from text and data mining (TDM) and AI training. It is not designed to serve as a licensing mechanism. This is by design, and for good reason.

### Opt-Out Is Not Licensing

Opt-out declarations, as supported by TDM·AI, are statements of reservation or permission. They are meant to communicate, in a standardised and interoperable way, that certain rights are (not) granted for specific AI-related uses of content. This applies broadly, including to unknown or future actors. It’s a public signal, not a bilateral agreement.

By contrast, licensing and opt-in mechanisms are contractual, negotiated, and typically confidential. A content license — whether exclusive or non-exclusive — is a transaction between identified parties. These transactions often involve terms and conditions that vary widely and are not publicly disclosed. For example, a publisher may license a dataset to one model developer under specific restrictions while denying access to others.

Because of their private, individualised nature, such agreements do not require a public vocabulary or machine-readable format. In practice, content licensing functions through legal contracts, not declarations or protocols.

### Different Purposes, Different Infrastructures

* Opt-out declarations serve a preventive purpose: They are meant to stop unauthorised or undesired use of content by signaling a refusal of permission under relevant exceptions (e.g., the DSM TDM exception).
* Opt-in agreements, on the other hand, are permissive and context-specific: They are about granting access under agreed terms.

### Exceptions vs. Licenses

TDM·AI is built around the EU legal framework, particularly the TDM exception in Article 4 of the Copyright in the Digital Single Market (CDSM) Directive – **but not limited to this legislation**. Under this framework:

* By default, AI model developers may use copyrighted works for TDM unless a rightsholder explicitly opts out in an appropriate manner, such as machine-readable means.
* An opt-out overrides the exception, restoring the need for explicit permission.

This is different from licensing, which would apply regardless of the exception. Opt-outs do not grant rights; they reserve them by revoking the exception.

In this context, TDM·AI allows rightsholders to signal their refusal of permission in a way that is compatible with EU law – and interoperable across the internet. Licensing, being a separate legal process, falls outside the scope of this technical method.

### Complementary, Not Redundant

A rightsholder may declare an opt-out using TDM·AI while still entering into licensing deals with specific parties. This is not a contradiction. In fact, it is expected:

* The opt-out sets a default position: no permission granted.
* A license creates an exception to that default: permission granted under specified terms.

TDM·AI makes no assumptions about licensing and does not interfere with it. It provides the public, standardised layer for opt-out declarations — nothing more, nothing less.

<br>


# Registries for Training Preferences

## A Standards-Based Model for Declaring Usage Restrictions

Creators and rightsholders who wish to express preferences about the use of their works in AI training can do so through a registry-based model for rights declarations. This model enables preferences to be expressed once and made discoverable to all relevant parties — including platforms, dataset curators, and AI model developers — in a machine-readable, standardised format.

The registry model relies on **federated content registries**: publicly accessible, verifiable directories that record metadata about how digital works may or may not be used, particularly in contexts such as text and data mining (TDM) and the training of AI models.

The **TDM·AI** protocol supports declarations that follow a shared structure, including:

* **ISCC codes** for uniquely identifying digital content, regardless of format or storage location
* **Opt-out categories** (e.g. for general TDM, AI training, generative AI training)
* **Machine-readable data formats** (e.g. JSON-LD) for integration into platforms, crawlers, and compliance tools

## Why Registries Are Needed

In the current digital environment, content is often collected and processed without direct interaction between the person using the content and the person who created it. Rights or usage permissions — if they exist at all — are typically buried in licensing terms, typically not machine-readable, hard to resolve at scale, or platform-specific.

Registries offer a structural solution. By linking declarations to the content itself (via ISCC fingerprints) and publishing them in a publicly accessible and verifiable form, they allow preferences to be made visible and actionable across systems.

This model provides:

* **Transparency** – Anyone can check whether a content asset is subject to a declared restriction
* **Consistency** – A shared vocabulary and format ensure that declarations are interpreted uniformly
* **Scalability** – The registry architecture supports high-volume indexing and look-up across billions of assets
* **Independence** – Registries are technology-neutral and not tied to a specific vendor, platform, or jurisdiction

Because declarations are made independently of the content’s location (e.g. not limited to a specific website or platform), they remain valid and resolvable even as content circulates across the internet or is reused in other contexts.


# Federated Registries Explained

## What Is a Federated Registry?

A federated registry is a network of independently operated directories that synchronise their entries to provide a **unified, globally searchable database** of content and metadata declarations. Rather than relying on a single central authority, multiple registries participate in a **shared protocol**, enabling:

* Distributed operation across sectors and jurisdictions
* Trust through digital signatures and public key infrastructure
* High-availability querying by AI developers, search engines, and compliance tools

Each registry publishes **minimal metadata** derived from the original declaration – just enough to confirm the rights status of a piece of content identified by an ISCC code.

## How It Works

Registries use a distributed data structure known as a **Distributed Hash Table (DHT)** to store and resolve declarations. Each registry entry is indexed using cryptographic identifiers, including:

* An **ISCC code** as a persistent fingerprint of the media asset
* A **Declaration (CID)** derived from metadata, hashes, and signer identity
* A **Decentralized Identifier (DID)** referencing the declaring party
* A **timestamp** with cryptographic signatures for authenticity and traceability
* Cryptographic **signatures** that allow to verify the integrity of the full metadata record

This system ensures that anyone – whether an online platform, AI developer, or rights organisation – can verify the declared rights associated with a specific work using public registry data. Verification does not require access to full metadata or private storage systems. The declaring party’s identity is referenced through a Decentralized Identifier (DID) and supported by qualified certificates or verifiable credentials, enabling trust without disclosing personal information.

## Operated by Many, Working as One

Registries can be run by creator groups, collecting societies, standards bodies, or compliance platforms. Each operator:

* Maintains its own node in the registry network
* Publishes declarations relevant to its members or mandate
* Can mirror or subscribe to declarations from other nodes

Because all registries speak the same "language" – a common metadata schema and synchronisation protocol – they function as **parts of a shared infrastructure**, even if operated independently.

## Tailored Access, Public Discovery

The registry layer is **public by design**—anyone can query it to check whether a content asset has been declared for a particular use case. At the same time, access to the full declaration data (stored separately) can be governed by access policies, ensuring **flexibility in governance and compliance**.


# Federated Registries vs. Centralised Approaches

As AI systems rely on public digital content for training, creators and rightsholders are looking for ways to express reservations about such uses. In response, a range of opt-out mechanisms have emerged – yet many raise concerns about their efficiency, accessibility, and robustness. Among these methods, centralised registries have been proposed to collect and store declarations, but such approaches have drawn criticism from creators and rightsholders for imposing costs and unnecessary burdens, overlooking sectoral diversity, and centralising control.

The Liccium registry infrastructure offers an alternative: a federated model of expressing AI-related preferences. It avoids central, monopolistic and expensive control, supports a wide range of sector-specific metadata, workflows and license models, and ensures declarations are verifiable, scalable, and accessible – without placing new burdens on creators.

This page explains the difference between centralised and federated registry models, how the Liccium system works, and why this approach addresses creator concerns.

## Concerns About Centralised Registries

From creators and publishers, several recurring concerns have been identified:

* Central control and governance risks: Who owns, controls and operates the registry?&#x20;
* Lack of sectoral alignment: Different sectors (books, press, photos) have their own specific metadata, workflows, and service providers.
* Duplicate effort: Centralised systems may require creators to manually register works that are already distributed via sectoral channels.
* Copyright re-affirmation: Creators fear being forced to re-prove or re-claim rights they hold by default.
* Language and region: Multilingual or jurisdiction-specific support is often lacking.
* Single point of failure: If a central database is taken offline or corrupted, access is lost.

## The Liccium Approach: Federated Registries

The Liccium registry infrastructure is designed to overcome these challenges through a federated, API-based, and peer-to-peer (p2p) model.

### **Core Principles**

* No central authority: Liccium registries operate on a p2p network – no single server, owner, or gatekeeper.
* Declaration = rights + signature: Each declaration includes cryptographic proof of identity and intent.
* Open API standard: All declarations follow a common specification (see [dev.liccium.com](https://dev.liccium.com/)).
* Sector integration: Any service provider (publisher, distributor, CMS, DAM, etc.) can implement the API and submit declarations on behalf of creators.
* Verifiable identity: Each declaring party is authenticated — via certificates or Verifiable Credentials.
* Scalable access: Declarations are synchronised across the network and accessible via a public registry index, not a centralised database.

### **What This Means in Practice**

* Creators are not burdened: Declarations can be made automatically via existing software (e.g. publishing tools, DAM systems, standard user apps and platforms). No need for creators to fill out separate registry forms.
* Avoidance of administration costs: Using a federated opt-out system avoids costs for redundant, duplicate data transmission where only the expression of a rights reservation is intended.
* Rightsholders maintain data sovereignty: Databases remain closer to the respective data sovereigns (such as publishers, public authorities, museums, libraries and CMOs).
* No need to re-affirm copyright: The Liccium model is not about establishing or proving ownership or copyright – it's about enabling the easy, secure and effective expression of preferences with regards to the exception (TDM), verifiably and at scale.
* Regional and sector-specific support: Any organisation – including creator or membership organisations, CMOs, or commercial service providers – can run a registry node or integrate the API into their tools. This also allows development of licensing models and effective bundling of licensing offers.
* Extensible metadata: Sector-specific metadata (e.g. IPTC for photos, ONIX for books) can be included or referenced. The common opt-out vocabulary (as defined here on [tdmai.org](https://tdmai.org/), based in the IETF vocabulary) is minimal and non-intrusive. Redefinition of established industry standards is unnecessary.
* Public auditability and resilience: All declarations are publicly resolvable, tamper-evident, and can be update over time – even if one registry node goes offline.

## Comparison Table

<table><thead><tr><th width="184.55859375">Feature</th><th width="182.66796875">Centralised Registry</th><th>Liccium Federated Registry</th></tr></thead><tbody><tr><td>Control model</td><td>Single operator</td><td>Decentralised, p2p</td></tr><tr><td>Governance</td><td>Central authority</td><td>Open protocol, sector participation</td></tr><tr><td>Authentication</td><td>Varies</td><td>Cryptographic identity (VCs, certificates)</td></tr><tr><td>Creator interaction</td><td>Direct submission required</td><td>Direct or indirect via trusted tools/providers</td></tr><tr><td>Metadata standards</td><td>Prescribed or inflexible</td><td>Flexible, sector-specific</td></tr><tr><td>Regional / language support</td><td>Limited</td><td>Local providers, multilingual</td></tr><tr><td>Copyright affirmation</td><td>Often required</td><td>Not required</td></tr><tr><td>Redundancy</td><td>Single point of failure</td><td>Distributed, resilient</td></tr><tr><td>Public discoverability</td><td>Controlled access</td><td>Open, verifiable index</td></tr><tr><td>Compliance support<br>(AI Act)</td><td>Varies</td><td>Full alignment with opt-out vocab</td></tr></tbody></table>

## Benefits of the Federated Model

In addition to solving the core concerns, the Liccium model provides:

* Interoperability across sectors, tools, and jurisdictions
* Compliance readiness for the EU AI Act and other frameworks
* Future-proof design for evolving vocabularies and metadata information (e.g. AI inference use cases)
* Trust infrastructure built on cryptographic signatures and DID/VC standards
* Integration with discovery: registry entries are publicly resolvable, inseparably bound to the content.

## Conclusion

The notion of a “registry” need not imply central control or burden on creators. The Liccium federated registry infrastructure enables decentralised, verifiable, and scalable declarations that are easy to perform and respect the diversity of content sectors – while providing reliable compliance signals for AI developers, platforms, and regulators.

By shifting from centralised control to distributed trust, Liccium offers a new foundation for rights reservation, creator agency, and cross-sector transparency in the AI age.

\ <br>


# TDMrep Vocabulary

2025-11-04

This section defines the **W3C TDMrep vocabulary** as used by the TDM·AI Protocol to express machine-readable rights reservations for digital content. Rightsholders and publishers use TDMrep declarations to signal their preferences regarding text and data mining (TDM) – including AI training and generative AI training – and publish those declarations to the TDM·AI registry, where they become persistently discoverable and verifiable.

{% hint style="info" %}
**Note on Alignment with IETF Drafts and Current Status**

The TDM·AI Protocol currently uses the **W3C TDMRep vocabulary** ([https://www.w3.org/community/reports/tdmrep](https://www.w3.org/community/reports/tdmrep/CG-FINAL-tdmrep-20240510/)) as the basis for expressing usage preferences. TDMrep is a W3C Community Group Final Report providing a standardised mechanism for rightsholders to declare TDM reservations in a machine-readable way.

We have chosen TDMrep as a stable and broadly adopted vocabulary that aligns well with the registry-based attachment model used by TDM·AI. Unlike domain- or location-based signaling, Liccium binds TDMrep declarations to individual digital assets via ISCC fingerprints (ISO 24138:2024), making preferences persistent and verifiable regardless of how content is distributed.
{% endhint %}

## Vocabulary Definition

The TDMrep vocabulary defines two terms:

| Category Name   | Key               | Description                                                                                                                                   |
| --------------- | ----------------- | --------------------------------------------------------------------------------------------------------------------------------------------- |
| TDM Reservation | `tdm-reservation` | Indicates whether TDM rights are reserved (`1`) or not reserved (`0`) for a resource                                                          |
| TDM Policy      | `tdm-policy`      | A URI pointing to a human- or machine-readable policy document specifying the conditions under which the content may be used for TDM purposes |

`tdm-reservation` is the required field. `tdm-policy` is optional and only meaningful when `tdm-reservation` is set to `1`.

### 1. TDM Reservation

`tdm-reservation` is a binary signal. A value of `1` means the rightsholder explicitly reserves all TDM rights – automated processing, including AI training, is not permitted without authorisation. A value of `0` means TDM use is permitted without contacting the rightsholder.

This single field covers the full scope of TDM activities, including text and data mining for commercial purposes, AI model training, and generative AI training.

### 2. TDM Policy

`tdm-policy` is an optional URI that links to a policy document where the rightsholder may specify more granular conditions – for example, permitted uses, licensing terms, or the legal basis for the reservation.

When present, automated systems are expected to retrieve and process the linked policy document to determine whether a specific use is permitted under the stated terms.

**Example:**

```json
{
  "tdm-reservation": 1,
  "tdm-policy": "https://example.com/tdm-policy"
}
```

A policy reference without a reservation has no effect. The `tdm-reservation` field must be set to `1` for the policy to be operative.

### Relationship to TDM·AI Registry Declarations

In the TDM·AI Protocol, TDMrep terms are not expressed at the domain level or embedded in content files. Instead, they are bound to ISCC fingerprints and published as signed declarations to federated registries. This registry-based approach ensures that rights reservations remain accessible even when content is shared, reformatted, or metadata is stripped from the original file.

Any party generating the ISCC for a piece of content can query the registry and retrieve the associated TDMrep declaration – without needing to know the original source domain or hosting location.


# JSON Format for TDMrep Declarations

2025-11-04

This section defines the JSON format used to express TDMrep declarations under the TDM·AI Protocol. TDMrep uses a simple two-field structure bound to an ISCC fingerprint, making it straightforward to implement while remaining fully machine-readable.

## Declaration Format

Each declaration consists of a flat key-value structure:

```json
{
  "iscc": "ISCC:EXAMPLE5QH7FTV7N5YVD5UMF4TUKFFGDGCOI4UDFKE4FNPW6C3L7J2Y",
  "tdm-reservation": 1,
  "tdm-policy": "https://example.com/tdm-policy"
}
```

### **Keys**

* `tdm-reservation` – indicates whether TDM rights are reserved. Required.
* `tdm-policy` – a URI pointing to a policy document specifying the conditions under which content may be used for TDM. Optional; only meaningful when `tdm-reservation` is `1`.

### **Values**

* `1` – TDM rights reserved; automated processing including AI training is not permitted without authorisation
* `0` – TDM rights not reserved; use is permitted without contacting the rightsholder

## Required and Optional Fields

Each declaration **MUST** include:

* `iscc` – a content-derived identifier based on ISO 24138:2024, representing the specific asset to which the declaration applies.
* `tdm-reservation` – the reservation flag (`1` or `0`).

Each declaration **MAY** include:

* `tdm-policy` – a URI to a human- or machine-readable policy document.

## Example Use Cases

### Example 1: Full TDM Reservation

All TDM rights reserved. No automated processing or AI training is permitted without authorisation.

```json
{
  "iscc": "ISCC:EXAMPLE5QH7FTV7N5YVD5UMF4TUKFFGDGCOI4UDFKE4FNPW6C3L7J2Y",
  "tdm-reservation": 1
}
```

#### Example 2: TDM Reservation with Policy Reference

Rights are reserved and a policy document specifies the conditions under which licensing may be available.

```json
{
  "iscc": "ISCC:EXAMPLE5QH7FTV7N5YVD5UMF4TUKFFGDGCOI4UDFKE4FNPW6C3L7J2Y",
  "tdm-reservation": 1,
  "tdm-policy": "https://example.com/tdm-policy"
}
```

#### Example 3: No Reservation

TDM use is explicitly permitted without contacting the rightsholder.

```json
{
  "iscc": "ISCC:EXAMPLE5QH7FTV7N5YVD5UMF4TUKFFGDGCOI4UDFKE4FNPW6C3L7J2Y",
  "tdm-reservation": 0
}
```


# Semantic Model

2025-04-11

Liccium declarations are designed to be both machine-readable and semantically interoperable. Each declaration references a versioned JSON Schema for structural validation, a JSON-LD context for semantic interpretation, and is grounded in a formal OWL ontology. Together these three artefacts define how TDMrep declarations can be understood, validated, and processed consistently across systems.

### Schema

The Liccium JSON Schema provides the formal structure for validating declaration payloads. It defines required and optional fields, permitted value ranges, and the overall structure of a declaration including its plugin metadata.

```
https://w3id.org/liccium/schema/0.3.0.json
```

Every declaration includes a `$schema` reference to this URL. Implementations should validate declarations against it before processing or honouring them.

Within the schema, TDMrep preferences are expressed under `liccium_plugins.tdmrep`:

```json
{
  "tdm-reservation": {
    "type": "integer",
    "enum": [0, 1],
    "description": "1 = TDM rights reserved; 0 = TDM rights not reserved."
  },
  "tdm-policy": {
    "type": "string",
    "format": "uri",
    "description": "URL linking to a TDM policy document."
  }
}
```

### JSON-LD Context

The Liccium JSON-LD context maps declaration fields to their formal semantic identifiers, allowing linked data agents to interpret declarations consistently.

```
https://w3id.org/liccium/context/0.3.0.json
```

Each declaration includes a `@context` reference to this URL. The context maps TDMrep terms to their canonical identifiers under the W3C TDMrep namespace (`https://www.w3.org/ns/tdmrep#`) and the Liccium ontology namespace (`https://w3id.org/liccium/ont/`).

### Ontology

The Liccium OWL ontology formally defines the properties used in declarations as RDF data properties, including their domain, range, and human-readable descriptions.

```
https://w3id.org/liccium/context/0.3.0.json
```

The two TDMrep properties are defined as follows.

`tdm-reservation` is a datatype property with range `xsd:integer`, accepting values `0` or `1`. A value of `1` means TDM rights are reserved and automated processing including AI training is not permitted without authorisation from the rightsholder. A value of `0` means TDM rights are not reserved and use is permitted without contacting the rightsholder.

`tdm-policy` is a datatype property with range `xsd:anyURI`. It links to a policy document – human- or machine-readable – specifying the conditions under which content may be used for TDM purposes. It is only meaningful when `tdm-reservation` is set to `1`.

Both properties have domain `liccium:Declaration`, meaning they are properties of a signed declaration bound to an ISCC fingerprint – not properties of the content file itself.


# Usage Preferences Vocabulary

2025-11-04

This section defines the controlled vocabulary used to express AI preferences under the **TDM·AI Protocol**. The vocabulary enables machine-readable communication of preferences regarding the use of digital content for automated processing activities, including AI training and generative AI training.

{% hint style="info" %}
**Note on Alignment with IETF Drafts and Current Status**

This specification on the Usage Preferences Vocabulary for the TDM·AI Protocol currently aligns with version 02 of the draft‑ietf‑aipref‑vocab‑02 published by the Internet Engineering Task Force (IETF) as of July 21 2025. [datatracker.ietf.org+2datatracker.ietf.org+2](https://datatracker.ietf.org/doc/html/draft-ietf-aipref-vocab-02?utm_source=chatgpt.com)\
\
It is important to note that the IETF work on the “AI Preferences” vocabulary remains a **work in progress**. The draft is still under discussion, subject to change, and does not yet constitute a finalized standard. [datatracker.ietf.org+1](https://datatracker.ietf.org/doc/draft-ietf-aipref-vocab/02/?utm_source=chatgpt.com)\
\
We have chosen to use this particular draft version as a **reference point** to illustrate how a domain-based attachment mechanism (as proposed by the IETF) can be **translated** into a registry-based system — that is, how preference declarations can be persistently and verifiably associated with individual digital assets via a registry rather than relying solely on domain- or location-based signaling.\
\
As the IETF draft evolves, we expect to revisit and update our alignment accordingly. Until then, this documentation should be regarded as **preliminary guidance**, not a definitive implementation.
{% endhint %}

| Category Name                         | Key           | Description                                                       |
| ------------------------------------- | ------------- | ----------------------------------------------------------------- |
| Automated Processing (formerly `tdm`) | `all`         | Top-level category covering all automated analysis and processing |
| AI Training                           | `train-ai`    | General-purpose or task-specific training of AI models            |
| Generative AI Training                | `train-genai` | Training models to produce synthetic content                      |
| AI Use (Inference)                    | `ai-use`      | Using assets as inputs to operate trained AI models               |
| Search                                | `search`      | Using assets in search engines or discovery applications          |

Each category may be declared independently. However, restrictions follow a strict hierarchy: opting out of a higher-level category (e.g., `all`) implies restriction of all its subordinate categories (`train-ai`, `train-genai`, `ai-use`, and `search`).

## Vocabulary Definition <a href="#name-vocabulary-definition" id="name-vocabulary-definition"></a>

This section defines the categories of use in the vocabulary, quoted from the IETF , <https://www.ietf.org/archive/id/draft-ietf-aipref-vocab-02.html#section-4>.

### 1. Automated Processing Category <a href="#name-automated-processing-catego" id="name-automated-processing-catego"></a>

"The act of using one or more assets in the context of automated processing aimed at analysing text and data in order to generate information which includes but is not limited to patterns, trends and correlations.

The use of assets for automated processing encompasses all the subsequent categories."

### 2. AI Training Category <a href="#name-ai-training-category" id="name-ai-training-category"></a>

"The act of training machine learning models or artificial intelligence (AI).

The use of assets for AI Training is a proper subset of Automated Processing usage"

### 3. Generative AI Training Category <a href="#name-generative-ai-training-cate" id="name-generative-ai-training-cate"></a>

"The act of training general purpose AI models that have the capacity to generate text, images or other forms of synthetic content, or the act of training more specialised AI models that have the purpose of generating text, images or other forms of synthetic content.

The use of assets for Generative AI Training is a proper subset of AI Training usage."

### 4. AI Use Category

"The act of using one or more assets as input to a trained AI/ML model as part of the operation of that model (as opposed to the training of the model).

The use of assets for AI Use is a proper subset of Automated Processing usage."

### 5. Search Category

"Using one or more assets in a search application that directs users to the location from which the assets were retrieved."

The purpose of defining a distinct Search category is to allow preferences to be expressed about search applications, independent of other categories of use. A distinct Search category allows for preferences specific to search applications, even if the use of AI is involved in their implementation.

The use of assets for Search is a proper subset of Automated Processing usage."


# JSON Format for Usage Declarations

2025-11-04

{% hint style="info" %}
**Note on Alignment with IETF Drafts and Current Status**

This specification on the Usage Preferences Vocabulary for the TDM·AI Protocol currently aligns with version 02 of the draft‑ietf‑aipref‑vocab‑02 published by the Internet Engineering Task Force (IETF) as of July 21 2025. [datatracker.ietf.org+2datatracker.ietf.org+2](https://datatracker.ietf.org/doc/html/draft-ietf-aipref-vocab-02?utm_source=chatgpt.com)\
\
It is important to note that the IETF work on the “AI Preferences” vocabulary remains a **work in progress**. The draft is still under discussion, subject to change, and does not yet constitute a finalized standard. [datatracker.ietf.org+1](https://datatracker.ietf.org/doc/draft-ietf-aipref-vocab/02/?utm_source=chatgpt.com)\
\
We have chosen to use this particular draft version as a **reference point** to illustrate how a domain-based attachment mechanism (as proposed by the IETF) can be **translated** into a registry-based system — that is, how preference declarations can be persistently and verifiably associated with individual digital assets via a registry rather than relying solely on domain- or location-based signaling.\
\
As the IETF draft evolves, we expect to revisit and update our alignment accordingly. Until then, this documentation should be regarded as **preliminary guidance**, not a definitive implementation.
{% endhint %}

This section defines the JSON format used to express opt-out and permission declarations under the **TDM·AI Protocol**. The format enables machine-readable communication of rights regarding the use of digital content for:

* Automated Processing (`all`)
* AI Training (`train-ai`)
* Generative AI Training (`train-genai`)
* AI Use / Inference (`ai-use`)
* Search (`search`)

The structure follows the [IETF AI Preferences Vocabulary](https://www.ietf.org/archive/id/draft-ietf-aipref-vocab-02.html) and [Attachment Mechanisms](https://www.ietf.org/archive/id/draft-ietf-aipref-attach-02.html), using compact single-byte tokens defined in Section 3.3.4:

* `true` — explicitly allowed
* `false` — explicitly disallowed

## Declaration Format

Each declaration consists of a flat key-value structure where usage types are associated with a value:

```json
{
  "all": "false",
  "ai-use": "true",
  "search": "true"
}
```

In this example:

* Automated processing and AI-training and Generative AI Training are disallowed
* AI inference and search are permitted

**Keys**

* `all` – General automated processing (formerly TDM)
* `train-ai` – AI model training
* `train-genai` – Generative model training
* `ai-use` – Use of content as input to a deployed AI model (inference)
* `search` – Use in search applications

**Values**

* `true` — allowed
* `false` — disallowed

## Inheritance and Overrides

The keys `all`, `train-ai`, and `train-genai` are hierarchically nested:

* `train-ai` inherits from `all`
* `train-genai` inherits from `train-ai`

Subordinate values inherit from parent values unless explicitly overridden. \
Other keys (`ai-use`, `search`) are independent and must be declared directly.

*Example:*

```json
{
  "all": "true",
  "train-genai": "false"
}
```

→ All processing is allowed except generative AI training, which is explicitly disallowed.

## Required and Optional Fields

Each declaration **MUST** include:

* `iscc`: A unique content identifier based on the ISCC standard, representing the specific asset to which the usage preferences apply.

Each declaration **MAY** include the following optional top-level fields:

* `intent`: A short code indicating the purpose of the declaration. Valid values are:
  * `activate` — for initial registration (default if omitted)
  * `update` — for a revision or change by the original declaring party
  * `supercede` — for an automatic update triggered by dependency or policy changes
* `summary`: A brief human-readable explanation of the declared usage rights, such as the scope of opt-out or permission. (optional)
* `policy`: A human-readable reference to relevant legal or regulatory frameworks that inform the declaration, e.g., EU AI Act, CDSM Directive, national copyright laws. (optional)

These optional fields provide additional context but do not affect the machine-readable enforcement of the core usage preferences.

## Example Use Cases for Opt-Out Declarations

### Example 1: Full Reservation of Automated Processing (TDM)

This declaration opts out of all usage categories—automated processing, AI training, and generative AI training – by setting the top-level category (`all`) to `n`. No overrides are needed for subordinate categories.

{% code overflow="wrap" %}

```json
{
  "iscc": "ISCC:EXAMPLE5QH7FTV7N5YVD5UMF4TUKFFGDGCOI4UDFKE4FNPW6C3L7J2Y",
  "all": "false",
  "summary": "Content must not be used for text and data mining, AI training, or generative AI training.",
  "policy": "The use of this work for text and data mining (TDM) is not permitted. This includes any automated analytical technique aimed at analyzing text or data in digital form to generate information, such as patterns, trends, or correlations. As a result, the work may also not be used for training general-purpose AI models or other systems, including those designed to generate synthetic content. This reservation is made in accordance with Article 4(3) of Directive 2019/790 (CDSM Directive)."
}
```

{% endcode %}

This compact structure uses the `all` category to disallow all subordinate types of use (`train-ai`, `train-genai`) through hierarchical inheritance.

### Example 2: AI Training Reserved

This declaration permits general automated processing (e.g. indexing, or non-AI analytical uses) while explicitly disallowing AI training. Since `train-genai` is a subset of `train-ai`, it is also implicitly disallowed and does not need to be stated separately.

{% code overflow="wrap" %}

```json
{
  "iscc": "ISCC:EXAMPLE5QH7FTV7N5YVD5UMF4TUKFFGDGCOI4UDFKE4FNPW6C3L7J2Y",
  "all": "true",
  "train-ai": "false",
  "summary": "Content may be used for text and data mining but must not be used for AI training or generative AI training.",
"policy": "The use of this work to train AI models is not permitted. This includes training general-purpose AI systems or other models capable of performing a wide range of tasks such as labeling, classification, pattern recognition, decision-making, or semantic content understanding. Use of the work for training generative AI models is also prohibited. However, text and data mining (TDM) is permitted in accordance with Article 4 of Directive 2019/790 (CDSM Directive), provided it does not serve the purpose of model training."
}
```

{% endcode %}

### Example 3: Generative AI Training Reserved

This declaration allows general automated processing and AI training for non-generative purposes, while explicitly disallowing generative AI training by overriding the inherited permission.

{% code overflow="wrap" %}

```json
{
  "iscc": "ISCC:EXAMPLE7UXMJCB6AVW4UHYMGYF6NNDPZKHQWQK5ZYPQJNPZAKGMYZQ",
  "all": "true",
  "train-genai": "false",
  "summary": "Content may be used for TDM and for training non-generative AI models, but not for generative AI training.",
  "policy": "The use of this work to train AI models that are either (a) general-purpose AI systems with the capacity to generate synthetic content such as text, images, audio, or video, or (b) other types of AI systems whose primary purpose is the generation of such content, is not permitted. Text and Data Mining (TDM) is allowed for non-generative purposes, including training AI systems that do not produce synthetic outputs, in accordance with Article 4 of Directive 2019/790 (CDSM Directive), and for scientific research or temporary reproduction under Article 5(1) of Directive 2001/29/EC."
}
```

{% endcode %}

### Example 4: Full Reservation Including Inference and Search

This declaration reserves all forms of automated processing, AI training, generative training, **and** downstream usage for inference and search.

{% code overflow="wrap" %}

```json
{
  "iscc": "ISCC:EXAMPLE99XXF42U8AV4EAVFJ6CE2ZH6MPAKMVDKAP5WZRE7YZU2U4FC",
  "all": "false",
  "ai-use": "false",
  "search": "false",
  "summary": "Content must not be used for any form of automated processing, AI training, inference, or search."
}
```

{% endcode %}

In this example, `train-ai` and `train-genai` are implicitly restricted by `all: "false"`, and `ai-use` and `search` are explicitly disallowed.

### Example 5: Inference and Search Permitted, Generative Training Reserved

This declaration permits general processing, non-generative AI training, inference, and search—but disallows generative AI training.

{% code overflow="wrap" %}

```json
{
  "iscc": "ISCC:EXAMPLE1YVP4YBZPXVFXMBTBKXGPV5VF6A7JHYYK5R45MEJJSZZDRQU",
  "all": "true",
  "train-genai": "false",
  "ai-use": "true",
  "search": "true",
  "summary": "Content may be used for search, inference, and training of non-generative AI systems, but not for generative AI training."
}
```

{% endcode %}

In this case:

* `all: "true"` permits all uses by default.
* `train-genai: "false"` overrides the default for generative training.
* `ai-use` and `search` are explicitly permitted to ensure downstream use is clearly allowed.


# Updates

2025-07-21

{% hint style="info" %}
**Note on Alignment with IETF Drafts and Current Status**

This specification on the Usage Preferences Vocabulary for the TDM·AI Protocol currently aligns with version 02 of the draft‑ietf‑aipref‑vocab‑02 published by the Internet Engineering Task Force (IETF) as of July 21 2025. [datatracker.ietf.org+2datatracker.ietf.org+2](https://datatracker.ietf.org/doc/html/draft-ietf-aipref-vocab-02?utm_source=chatgpt.com)\
\
It is important to note that the IETF work on the “AI Preferences” vocabulary remains a **work in progress**. The draft is still under discussion, subject to change, and does not yet constitute a finalized standard. [datatracker.ietf.org+1](https://datatracker.ietf.org/doc/draft-ietf-aipref-vocab/02/?utm_source=chatgpt.com)\
\
We have chosen to use this particular draft version as a **reference point** to illustrate how a domain-based attachment mechanism (as proposed by the IETF) can be **translated** into a registry-based system — that is, how preference declarations can be persistently and verifiably associated with individual digital assets via a registry rather than relying solely on domain- or location-based signaling.\
\
As the IETF draft evolves, we expect to revisit and update our alignment accordingly. Until then, this documentation should be regarded as **preliminary guidance**, not a definitive implementation.
{% endhint %}

## Updating Usage Preferences

A rightsholder who has previously declared a usage preference – whether to allow or disallow certain types of use – may later choose to **update** that declaration. An update is a machine-readable statement that supersedes a prior declaration for the same asset and explicitly communicates a change in permissions or restrictions.

While permission is generally assumed by default (in the absence of an opt-out or opt-in requirement), publishing an explicit update is important in scenarios where:

* A prior opt-out or usage restriction is already in circulation and may still be active in third-party systems
* The declaring party wishes to modify or replace an earlier declaration (e.g., due to licensing changes or policy shifts)
* There is a need to maintain continuity and transparency in the audit trail of usage preferences

## Declaration Format for Updates

Updates use the same JSON structure as any standard TDM·AI usage declaration. An update is intended to **replace** a previously issued declaration for the same asset.

To indicate that a declaration is an **update**, two optional fields may be included:

* `"intent": "update"` — signals that the purpose of this declaration is to revise or replace a prior statement.
* `"reference"` — provides the **Declaration ID** (a CID v1 hash) of the earlier declaration being replaced.

While `"intent"` serves as a **semantic hint** to systems, `"reference"` creates an **explicit link** to the prior declaration. This improves auditability and allows registries to track the update chain over time.

## Example: Update of a Previous Usage Reservation

This example lifts a prior reservation and re-authorises the use of the content for all types of automated processing, including TDM, AI training, and generative AI training.

{% code overflow="wrap" %}

```json
{
  "iscc": "ISCC:EXAMPLE5QH7FTV7N5YVD5UMF4TUKFFGDGCOI4UDFKE4FNPW6C3L7J2Y",
  "all": "true",
  "intent": "update",
  "reference": "bafyreibxxxxxxxyyyyyyzzzzzzzddeeeeeeeeeeeeeeeeeeeeeeeee"
}
```

{% endcode %}

In this case:

* `all: "true"` allows all uses (implicitly permitting `train-ai` and `train-genai`)
* `reference` points to the **Declaration ID** (CID v1) of the earlier declaration being replaced
* `intent: "update"` marks the purpose of the change and makes the revocation explicit and traceable for downstream systems.

Registries and consumers can use this linkage to deprecate or replace the earlier declaration in their compliance pipelines.

{% hint style="info" %}

### Why issue a revocation in the context of the EU CDSM?

Under the EU’s Copyright in the Digital Single Market Directive (CDSM Directive, 2019/790), the use of copyrighted content for Text and Data Mining (TDM) is governed by statutory exceptions rather than by traditional licensing.

Two Key Legal Provisions Define the Framework:

* Article 3 — Allows TDM for research organisations and cultural heritage institutions without needing prior permission. No opt-out is possible.
* Article 4 — Allows TDM for all other users, including commercial actors, unless the rightsholder has explicitly reserved their rights in a machine-readable way.

In practice, this means that:

> **If no reservation is declared under Article 4, TDM is permitted by default.**

This legal default – referred to in Article 4(3) – makes it lawful to mine content for patterns, trends, or correlations unless the rightsholder has clearly expressed a reservation.

Even though permission is assumed in the absence of a reservation, an explicit revocation of a previous reservation is often necessary or useful. This is because:

* Third-party systems may cache or store earlier declarations, continuing to enforce restrictions that are no longer intended.
* Archived opt-outs may persist in datasets used by AI developers or data brokers.
* Researchers and developers may require proof of permission, especially if access was previously denied.

A revocation declaration serves to:

* &#x20;Remove any lingering restrictions from prior machine-readable opt-outs
* Clearly signal that the rightsholder now allows TDM, AI training, or generative AI training

Revocations ensure transparency, help maintain up-to-date compliance, and support interoperability across evolving systems.
{% endhint %}


# JSON Schema Definition

2025-04-11

{% hint style="info" %}
**Note on Alignment with IETF Drafts and Current Status**

This specification on the Usage Preferences Vocabulary for the TDM·AI Protocol currently aligns with version 02 of the draft‑ietf‑aipref‑vocab‑02 published by the Internet Engineering Task Force (IETF) as of July 21 2025. [datatracker.ietf.org+2datatracker.ietf.org+2](https://datatracker.ietf.org/doc/html/draft-ietf-aipref-vocab-02?utm_source=chatgpt.com)\
\
It is important to note that the IETF work on the “AI Preferences” vocabulary remains a **work in progress**. The draft is still under discussion, subject to change, and does not yet constitute a finalized standard. [datatracker.ietf.org+1](https://datatracker.ietf.org/doc/draft-ietf-aipref-vocab/02/?utm_source=chatgpt.com)\
\
We have chosen to use this particular draft version as a **reference point** to illustrate how a domain-based attachment mechanism (as proposed by the IETF) can be **translated** into a registry-based system — that is, how preference declarations can be persistently and verifiably associated with individual digital assets via a registry rather than relying solely on domain- or location-based signaling.\
\
As the IETF draft evolves, we expect to revisit and update our alignment accordingly. Until then, this documentation should be regarded as **preliminary guidance**, not a definitive implementation.
{% endhint %}

This section provides the formal JSON Schema for validating declarations under the **TDM·AI Protocol**. The schema ensures that all declarations follow a consistent, machine-readable structure, enabling reliable processing across tools, registries, and compliance systems.

Each declaration expresses the rightsholder’s preferences regarding the use of their content for:

* Automated Processing (`all`)
* AI Training (`train-ai`)
* Generative AI Training (`train-genai`)
* AI Use / Inference (`ai-use`)
* Search (`search`)

Usage preferences are encoded using the IETF-compatible tokens:

* `"true"` — use is explicitly allowed
* `"false"` — use is explicitly disallowed (opt-out)

The schema defines required fields, permitted values, and optional metadata for lifecycle management such as updates and policy declarations.

Implementations are expected to validate usage reservation declarations against this schema before processing or honouring them. In specific use-cases, they MAY additionally want to reject or ignore any declarations containing unknown fields, i.e. by adding `"additionalProperties": false` at the root level of the schema.

The following schema suffices to validate conformance to this core data model.

{% code overflow="wrap" %}

```json
{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "$id": "https://docs.tdmai.org/technical-specification/json-schema-definition",
  "title": "TDM·AI Usage Declaration",
  "type": "object",
  "required": ["version", "iscc", "intent"],
  "properties": {
    "version": {
      "type": "string",
      "description": "The version of the reservation declaration schema.",
      "const": "1.0"
    },
    "iscc": {
      "type": "string",
      "description": "A unique content identifier using the ISCC standard (ISO 24138:2024)."
    },
    "all": {
      "type": "string",
      "description": "Preference for general automated processing.",
      "enum": ["true", "false"]
    },
    "train-ai": {
      "type": "string",
      "description": "Preference for AI training (non-generative).",
      "enum": ["true", "false"]
    },
    "train-genai": {
      "type": "string",
      "description": "Preference for generative AI training.",
      "enum": ["true", "false"]
    },
    "ai-use": {
      "type": "string",
      "description": "Preference for inference-time use of the asset as input to a trained model.",
      "enum": ["true", "false"]
    },
    "search": {
      "type": "string",
      "description": "Preference for inclusion in search applications.",
      "enum": ["true", "false"]
    },
    "reference": {
      "type": "string",
      "description": "Reference to a prior declaration (CID) this one updates or supersedes."
    },
    "intent": {
      "type": "string",
      "description": "Type of declaration in its lifecycle.",
      "enum": ["activate", "update", "supersede"],
      "default": "activate"
    },
    "summary": {
      "type": "string",
      "description": "Optional human-readable summary of declared permissions or reservations."
    },
    "policy": {
      "type": "string",
      "description": "Optional human-readable reference to the targeted compliance regime (e.g. CDSM, AI Act)."
    }
  },
  "additionalProperties": false
}
```

{% endcode %}


# Legal Basis

## Disclaimer on Legal Applicability

This vocabulary is designed to provide a standardised, machine-readable means for rightsholders to communicate usage preferences concerning the use of protected content for automated processing or text and data mining (TDM), artificial intelligence (AI) training, generative AI training, RAG/inference, and search, involving AI.

This vocabulary operates in the context of ongoing legal debate regarding the scope and applicability of statutory TDM exceptions – particularly whether such exceptions, as provided for in EU and national copyright laws, extend to the training of generative AI systems. Current academic and legal discourse, including the work of Dornis and Stober (2024), indicates divergent interpretations and unresolved questions in this area (Dornis, T.W. & Stober, S. (2024). Urheberrecht und Training generativer KI-Modelle – Technologische und juristische Grundlagen, Recht und Digitalisierung, Nomos Verlag.[ Open Access version](https://www.nomos-elibrary.de/10.5771/9783748949558/urheberrecht-und-training-generativer-ki-modelle?page=1); also available at SSRN:[ https://ssrn.com/abstract=4946214](https://ssrn.com/abstract=4946214)). Currently, the debate focuses on the evolving interpretation of Article 4(3) of Directive (EU) 2019/790 (CDSM Directive) and the interpretation of obligations in Article 53(1)(c) AI Act. Similar uncertainty exists in U.S. law, as evidenced in ongoing federal litigation such as Bartz v. Anthropic, where courts declined early dismissal of copyright claims involving AI training data.

The inclusion of terms and categories in this vocabulary does not constitute a legal determination of whether any given use is permitted or prohibited under applicable law. Rather, it reflects the intention of rightsholders to express usage rights, e.g. reserve rights, to the fullest extent permitted by law and to provide clear signals to users and AI developers in light of legal uncertainty.

Implementers and users of this vocabulary are advised to seek legal counsel regarding the specific application of copyright exceptions and limitations in their jurisdiction. The use of this vocabulary does not substitute for legal advice nor does it imply endorsement of any particular legal interpretation. The use of this vocabulary cannot substitute for formal legal consultation or override court interpretations or authoritative legal standards.

## The EU Copyright Directive

[Article 4 of the EU Copyright Directive 2019/790 (DSM Directive)](https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32019L0790) requires EU Member States to provide an "exception \[...] of this Directive for reproductions and extractions of lawfully accessible works and other subject matter for the purposes of text and data mining", which is defined as "any automated analytical technique aimed at analysing text and data in digital form in order to generate information which includes but is not limited to patterns, trends and correlations".

The DSM Directive distinguishes between TDM for scientific purposes and for other purposes. Article 3 provides a provision for scientific purposes, while Article 4 regulates all other purposes. Although Article 4 does not explicitly mention any specific, non-scientific purposes of TDM, it is assumed that acts of training models of AI include acts of TDM as an essential element.&#x20;

The applicability of Article 4 CDSM to AI training is referenced in Article 53(1)(c) of the AI Act, which obliges providers to respect copyright and opt-out declarations. While this interpretation is operationalised under the AI Act, the European Parliament has called for clearer legal provisions or a dedicated exception to regulate generative AI training (Motion for a European Parliament Resolution on Copyright and generative artificial intelligence – opportunities and challenges, p. 12 (2025/2058(INI)), published June 26, 2025, <https://www.europarl.europa.eu/doceo/document/JURI-PR-775433_EN.pdf>)

Article 4 (3) CDSM further elaborates that "the exception or limitation provided for in paragraph shall apply on condition that the use of works and other subject matter referred to in that paragraph has not been expressly reserved by their rightholders in an appropriate manner, such as machine-readable means in the case of content made publicly available online."&#x20;

## International Applicability of the Relevant Provisions within the EU Copyright Directive and the AI Act

The AI Act and the obligation to respect TDM opt-out declarations raise complex questions about their territorial applicability on cross-border uses. If, for instance, an AI provider in the US trains an AI model, using servers and content available in the US, it is questionable whether this provider must respect TDM opt-out declarations. However, Recital 106 of the AI Act expresses the principle that no provider shall gain a competitive advantage by applying lower copyright standards outside the EU.&#x20;

The AI Act assumes that the provider must respect the declaration at least when marketing these models on the European Union Market, in order “to ensure a level playing field among providers of general-purpose AI models where no provider should be able to gain a competitive advantage in the Union market by applying lower copyright standards than those provided in the Union.” (Recital 106 of the European AI Act). Even if this assumption is only part of a (non-binding) recital, it expresses the general objective of the Act, confirmed by obligations affecting providers irrespective of their seat such as those contained in § 53 of the AI Act.

Traditionally, copyright laws are territorial in the sense that reproductions during AI model training are governed by the copyright laws of the country where the training occurs, this would primarily suggest the applicability of US law to the case. The EU Copyright Directive, including its opt-out provision for text and data mining (TDM), is European Union (EU) legislation. As such, it primarily applies to EU member states and the acts of using copyrighted works on EU territory.

However, the scope of the AI Act, referring to the copyright provisions of the EU Copyright Directive goes beyond this regulation. The AI Act imposes very specific compliance duties, namely, “through state-of the art technologies” to comply with opt-out declaration (Art. 53 (1) lit. c) AI Act), on any provider placing general-purpose AI models on the EU market, regardless of their location including the obligation to implement a copyright policy and state-of-the-art technologies to comply with opt-outs (Code of Practice, Copyright Chapter). This means even if the AI training occurs outside the EU, such model must comply with the AI Act when the model or an application based thereon is offered on the European Union market. The Act's goal is to ensure fair competition and protect individual rights within the territorial scope of EU law. Therefore, if an AI model is not developed according to the AI Act, it cannot be offered in the EU, regardless of whether the content used for training has been created by a rightsholder based in the EU or on US territory.

Alternatively, *Rosati* and other legal scholars suggest that the application of the AI Act could depend on the localisation of another act of use, namely, on whether the AI provider has crawled data from websites hosted in the EU. *Rosati* also suggests that “terms of service of AI model providers that seek to remove any liability potentially arising from the use of their models might turn out to be ineffective towards third-party rightholders or users of their services.” This approach is consistent with the specific requirements of the AI Act, such as automatic crawling and the obligation to comply with machine-readable rights reservations. This solution emphasizes that if content is hosted in the EU, AI providers accessing this content by way of crawling during or as a part of AI training processes should comply with applicable rules in the EU, even if the subsequent training takes place elsewhere.

In summary, while traditional copyright principles may in effect limit territorial applicability, the AI Act seeks to ensure compliance with EU standards for AI models offered within the EU, potentially extending its reach beyond traditional territorial boundaries.

## Applicability to Rightsholders based outside the European Union

Creators and rightsholders from the US or from other countries outside the EU may benefit from the EU Copyright Directives if their content is used within the EU. The Parliament has also proposed establishing a central EUIPO registry of opt-outs, which could support international rightsholders in signaling their preferences effectively. (JURI Report, paras 6–9 and Explanatory Statement; <https://www.europarl.europa.eu/doceo/document/JURI-PR-775433_EN.pdf>).

If the content of US rightsholders is affected by uses happening on EU territory and is lawfully accessible within the EU, AI companies are marketing their products in the EU and EU-based users are using it, these AI companies must comply with the EU Copyright Directive. Notably, U.S. courts have begun scrutinising whether AI providers owe a duty to respect copyright when using publicly accessible content for training, as in Bartz v. Anthropic.

Therefore, US and other third country rightsholders may want to consider the instruments provided for their protection by the EU Copyright Directive if they want to control the use of their content within the EU. They can implement the opt-out provision to prevent their content from being used for TDM by AI providers marketing their products in the EU.

## Implementation Considerations for Machine-Readable Opt-Out

US and other third country rightsholders can use machine-readable means, such as the International Standard Content Code (ISCC) or other EU standardised formats as discussed in the Code of Practice (AI Code of Practice, Copyright Chapter, Measure 1.3), to identify their works and bind opt-out reservations that express in an automated way that their works should not be used for TDM. This would help ensure that EU-based entities can respect their opt-out.

Practical Steps – US and other third country rightsholders should:

1. Assess if their content is accessible within the EU.
2. Implement a machine-readable opt-out if they wish to prevent TDM through content that is accessible and marketed within the EU.
3. Monitor compliance and take necessary actions if their opt-out is not respected.
4. Monitor relevant litigation (e.g. Bartz v. Anthropic) in the U.S. and other jurisdictions to track how courts assess copyright liability in AI training contexts.

In summary, the EU Copyright Directive's opt-out provision can be applicable if an AI model or the AI application is marketed within the European Union. Implementing a machine-readable opt-out, using the ISCC as an identifier that can help to point to rights and opt-out reservations, can provide an effective and enforceable instrument to US and other third country rightsholders to manage their content's use.

## Sources

Rosati E. Infringing AI: Liability for AI-Generated Outputs under International, EU, and UK Copyright Law. European Journal of Risk Regulation. 2025;16(2):603-627. doi:10.1017/err.2024.72  Rendas/Hartmann; From Brussels to Brasília: How the EU AI Act Could Inspire Brazil’s Generative AI Copyright Policy; GRUR Int. 2024, p. 495.

HiQ Labs, Inc. v LinkedIn Corp. 31 F.4th 1180, 1187 n. 3 (9th Cir. 2022).

See Bartz v. Anthropic PBC, No. 3:24-cv-00518, Dkt. 231 (N.D. Cal. May 6, 2024), where the court denied Anthropic’s motion to dismiss copyright claims stemming from the alleged ingestion of protected content into AI models, citing unresolved factual and legal questions concerning fair use.

<br>


# Our Position Paper

An early position paper was submitted to the IAB Workshop on AI-CONTROL (aicontrolws) and accepted. The workshop took place in Washington in September 2024. Below, you can download the published version as PDF-document. Be aware that this version is outdated by the course of events.&#x20;

## Event

IAB Workshop on AI-CONTROL:\
<https://datatracker.ietf.org/group/aicontrolws/about/>\
Workshop Dates: 19-20 September 2024

## Title

TDM·AI – Making Unit-Based Opt-out Declarations to Providers of Generative AI

## Authors

Sebastian Posth (M.A.), Founder and CEO of Liccium.com\
Sabine Richly, German Qualified Lawyer (LL.M., MBA)

## Abstract&#x20;

TDM·AI is a protocol for creators and rightsholders to inseparably bind their machine-readable preferences for text and data mining (TDM) to digital media assets, specifically tailored for training models and applications of generative AI. TDM·AI addresses the main problem in controlling AI crawlers, namely the problem of metadata binding, by proposing a reliable method of soft-binding restrictions or permissions to use content for training models of generative AI to content-derived identifiers. The TDM·AI protocol utilises the International Standard Content Code (ISCC), a new ISO standard for the identification of digital media content (ISO 24138:2024) and Creator Credentials, based on W3C recommendation for cryptographically verifiable credentials, to ensure that verifiable and machine-readable declarations include proper attribution of preferences and claims to the legitimate rightsholders. Although the protocol has its origins in the European DSM Directive on Copyright 2019/790, Article 4, it may in many cases also be applicable to content published by rightsholders outside the EU.

## Download

{% file src="/files/6dQ7QLdar17yC62Qrtiy" %}


