Guides

    AI watermarking and your deliverables

    Text produced with frontier AI is becoming machine-detectable, and the parties who can run that check are the ones you answer to. The mark is not the problem. Work that cannot be explained is.

    By Gabriel Tremblay

    Published 2026-08-23 · Last reviewed 2026-08-23

    The short answer

    Anthropic began embedding watermarks in Claude's text output on 2 August 2026; Google has marked Gemini text with SynthID since 2023. The mark is statistical — carried in word choice, not hidden characters — so it survives copy-paste and cannot be stripped by reformatting. It carries no identifying information. Detection tooling is announced but not yet public, and coverage is prospective. A mark shows that AI was involved, not how it was used; only a contemporaneous record shows that.

    What changed, and who can see it

    The situation, rather than the technology, is what matters first: text produced with frontier AI is becoming machine-detectable, and the parties positioned to run that check are the ones a consultant or regulated expert answers to — sponsors, agencies, HTA reviewers, and clients enforcing their own contracts. Anthropic began embedding watermarks in Claude's text output on 2 August 2026. Google has used SynthID on Gemini text since 2023. Other providers are implementing their own schemes on their own timelines.

    Two qualifications matter. (1) Detection tooling is not yet publicly available — Anthropic has said a detection API is coming, and implementation details are not settled. (2) Coverage is prospective, not retroactive: models launched on or after 2 August 2026 mark at launch, older models are not yet covered, and text generated before that date is unaffected. AI regulation and transparency requirements are going to become steadily stricter, and we need to be in front of that rather than behind it. Exposure begins now and grows from here; it does not reach backwards into work already delivered.

    It means work augmented with AI will now carry a mark showing that AI was involved. It does not show how much, or how. A consultant or regulated expert who used AI to tighten prose under close human review and one who outsourced an entire deliverable can produce the same signal. The watermark does not distinguish them. A record does.

    What actually gets marked, and how

    Text and files are two different mechanisms with different properties, and conflating them leads to bad decisions.

    Text outputImage and file output
    MechanismStatistical watermark — a bias applied to which candidate word the model selects at each generation step.C2PA content credentials — a signed provenance manifest attached to the file.
    Where the signal livesIn the word choices themselves. Nothing is appended; no hidden characters, no metadata inserted into the text.In file metadata, genuinely separate from the visible content.
    What it can carryNo identifying information. It cannot be traced to a person, an organization or a conversation.Provenance detail, because a manifest is metadata and can hold it.
    Applies toGenerated text, including text pasted into another document.File outputs such as .png, .jpg and .svg. Not plain text.

    Because the text signal is carried in word choice, it survives copy-paste into another document and can persist through some editing. Provider-by-provider coverage is uneven and still moving, so the practical assumption is that any frontier-model text may be marked, or soon will be.

    What reformatting does not do

    You cannot strip a text watermark by cleaning up the file. Retyping the document, find-and-replace, normalising Unicode, converting between formats, pasting through a plain-text editor — none of it touches a signal carried in which words were chosen. There are no invisible characters to remove because none were inserted. Anyone selling a "cleaner" for AI text is selling something that does not do what it claims.

    Can it be removed, legally?

    As of August 2026, no general US or Canadian statute makes removing an AI watermark an offence on its own. Here is what it actually means, and what it actually breaks:

    • It breaches the provider's terms of use. The practical consequences are throttling, suspension or termination of access.
    • California's AI Transparency Act builds revocation into the supply chain. A covered provider that learns a licensee has stripped or disabled a latent disclosure must revoke that licensee's access within 96 hours. The consequence arrives through your vendor, not through a regulator.
    • In the EU, Article 50 places the marking duty on providers. The emerging reading is that deployers must not suppress or defeat provider-embedded disclosures. Where a deployer carries its own disclosure duty, removal works directly against it.
    • Deliberate removal is evidence of intent to conceal. This is the decisive point for regulated work. It converts a compliance question into a bad-faith question, in front of a regulator, an HTA agency, a government body, or a client that simply wants transparency — and bad faith is a far worse position to argue from than an unauthorized-tool finding.

    Article 50 of the EU AI Act, which requires providers to ensure generative output is marked in a machine-readable format and detectable as artificially generated, applies from 2 August 2026. That regulation is the driver behind the whole rollout, and it is moving in one direction.

    A watermark is not the problem

    The instinct is to treat detectability as a threat. It is only a threat to work that cannot be explained. Hiding it, or doing it secretly, is not the answer — transparency is.

    Start from what the field already accepts. AI-assisted and AI-augmented documents are legitimate. Guidance in HEOR assumes AI use and asks for human oversight, transparency about AI involvement, and traceability that survives scrutiny — not abstinence. ISPOR named AI the top HEOR trend for 2026–2027, and an ISPOR Working Group has published a taxonomy of generative AI applications across the HEOR workflow (Fleurence et al., Value in Health 2025;28(11):1601–1610). The same working group's report on health technology assessment is explicit that these models are to augment human work, not replace it — foundation models, it says, should augment “human activities for which humans remain fully accountable” (Fleurence et al., Value in Health 2025;28(2):175–183). It also calls for reporting guidance and checklists covering validation, reproducibility and bias. Nobody in that conversation is arguing that using AI is the violation — the argument is about the conditions, and every one of them is a condition about the record.

    What creates exposure is something else: unauthorized disclosure of confidential values, secrets and unpublished data, and the absence of any record of what was done. Those are two separable problems, and they have separate solutions. Cloaking addresses the first — the confidential value is replaced on your own computer, so the model never receives it. The audit trail — such as the Untraceable Protocol — addresses the second. Solve both and a mark on the deliverable is not a finding. It is corroboration: evidence consistent with a workflow you can already document.

    Why an AI humanizer has no place in regulated work

    A whole category of tools now offers to rewrite model output so it "reads as human" and clears AI detectors. For a consultant or regulated expert, adopting one is a bad trade on every axis that matters.

    • It does not do what it claims on text. The provider watermark is statistical, carried in word choice. A tool that paraphrases is, at best, guessing at which words carried the signal — and you have no way to verify the result, because the detection tooling is not public yet.
    • It breaches the provider's terms of use. The same exposure as stripping a mark directly: throttling, suspension or termination of access.
    • It puts you inside California's revocation mechanism. A covered provider that learns a licensee has disabled a latent disclosure must revoke that licensee's access within 96 hours. Using a humanizer is a way of arriving at that outcome indirectly.
    • It works against the EU Article 50 direction of travel. The emerging reading is that deployers must not suppress or defeat provider-embedded disclosures. Where you carry your own disclosure duty, a humanizer defeats it.
    • It is the intent problem in its clearest form. Stripping a mark can at least be argued as a formatting accident. Buying a tool whose stated purpose is to evade AI detection cannot. In front of a regulator, an HTA agency, a government body or a client that simply wants transparency, that purchase is the finding.
    • It degrades the deliverable. Paraphrasing for evasion optimises for reading as human, not for being correct. In regulated writing, where a number or a qualifier carries the meaning, that is the last thing you want a machine doing unsupervised.

    The tell is that a humanizer solves a problem you should not have. If the AI use was authorized, scoped and free of confidential values, there is nothing to evade. If it was not, the humanizer does not fix that — it just removes your ability to explain it.

    Our position, plainly. Untraceable does not hide, disguise, strip or degrade watermarks, and we would decline to build that. We treat them as one more element of provenance. The reason is straightforward: a company whose product is auditability cannot also sell concealment, and any vendor offering both is telling you what it actually values.

    Why the audit trail turns this to your advantage

    Set the two situations side by side. A sponsor detects a mark in a delivered dossier and asks about it.

    Without a recordWith the Untraceable audit
    When the answer is producedAfter the question, under time pressure.Before the question. The record predates it.
    What you can showRecollection and inference. Every answer sounds constructed, because it is being constructed.Which model, which version, which parameters, what instructions were sent, what came back, which cloaks were in play, what the cloak-coverage test, sentinel and breach test returned before transmission, who certified the run, and what the human reviewed and changed — alongside the risk-assessment documentation, the submission certification document and the privacy-assessment documents, provided with e-signature.
    Can it be handed over?Not without review — you do not know what it would reveal.Yes. Confidential values appear nowhere in it, so it goes to the sponsor without redaction.
    Who controls the conversationThe client, deciding whether an NDA was breached.You, because you can produce evidence on demand.

    The shareable audit is safe to give to a sponsor or an agency, while the complete version — the one including real values — is regenerated client-side and never leaves your own tenant. You are not choosing between transparency and confidentiality; the record is built so that you get both.

    There is a forward-looking argument too, and it does not need hype to land. Marking is rolling out unevenly and the regulation behind it is still moving. A firm that has been logging provenance all along is ready for a disclosure requirement whenever one arrives. A firm that has not will be reconstructing history under time pressure, which is the worst moment to discover what was never written down. The Untraceable Protocol is the specification for that record, published openly so any tool or team can conform to it.

    What to do now

    • Assume anything you produce with frontier AI from here is markable. Plan for visibility rather than against it — the planning-against options do not work and carry their own cost.
    • Read what your MSAs actually restrict. Most restrict disclosure of confidential information rather than AI use as such. That distinction determines what you genuinely need to fix, and it is often narrower than people assume. Does using AI break your NDA?
    • Do not adopt tools that promise undetectable output. They do not work on text watermarks, and adopting one creates the intent problem described above.
    • Keep confidential values out of the model. This is the disclosure question, and it is entirely independent of detectability — worth solving on its own terms. De-identification is not confidentiality
    • Keep a contemporaneous record of every AI submission. Reconstructed after the fact, it is worth much less — if anything — and will involve speculating and extrapolating the truth. AI for HEOR: what you can and can't put in a prompt
    • If you sponsor consulting work, say what you will permit. Silence produces shadow use, and shadow use produces exactly the exposure the silence was meant to avoid.

    The record is the whole answer here. If you want to see what it looks like, the Untraceable Protocol sets out what it must contain, the security architecture covers how the confidential values stay out of it, and a demo runs it on your own workflow.

    Frequently asked questions

    Can a sponsor tell that my deliverable was written with AI?

    Increasingly, yes. Anthropic began embedding a watermark in Claude's text output on 2 August 2026 and Google has marked Gemini text with SynthID since 2023, with other providers rolling out their own schemes. Detection tooling is not publicly available yet, and coverage is prospective rather than retroactive — but the direction is settled. Plan for your work to be visible rather than assume it will not be.

    Does a detected watermark prove the AI wrote my document?

    No. Detection is probabilistic and asymmetric: a mark indicates the text may have been processed by that model, not that the model authored it. Someone who used AI to translate or proofread their own original writing produces marked output from genuinely their own work. The reverse also holds — the absence of a mark is not evidence of human authorship.

    Can I remove the watermark by reformatting or retyping the document?

    No. The text watermark is statistical — a bias in which candidate word the model selects — so the signal lives in the word choices themselves. Nothing is appended and there are no hidden characters to strip. Retyping, find-and-replace, Unicode normalisation, format conversion and pasting through a plain-text editor do not touch it. Any tool sold as a watermark cleaner does not work on text.

    Is it illegal to remove an AI watermark?

    As of August 2026 no general US or Canadian statute makes removal a criminal or civil offence on its own — but that is the weakest reason not to do it. It breaches the provider's terms of use, and California's AI Transparency Act requires a covered provider that learns a licensee stripped a latent disclosure to revoke that licensee's access within 96 hours. For regulated work the decisive point is simpler: deliberate removal is evidence of intent to conceal, in front of a regulator, an HTA agency, a government body, or a client that simply wants transparency.

    Does Untraceable remove or suppress watermarks?

    No, and we would decline to build it. Cloaking governs what the model receives, not what it emits — the text that comes back is watermarked like any other model output. We treat a watermark as one more element of provenance, and pair it with an audit trail that shows the use was authorized, scoped and free of confidential values.

    What should I do about this now?

    Assume anything you produce with frontier AI from here is markable, keep confidential values out of the model, and keep a contemporaneous record of every AI submission. Check what your MSAs actually restrict — most govern disclosure of confidential information rather than AI use as such. A record made at the time is worth far more than one reconstructed under questioning.

    This guide is general information about how AI tools interact with confidentiality obligations. It is not legal advice, and it does not create any professional relationship. Confidentiality agreements vary — review your own agreements with qualified counsel before relying on any framework described here.

    Work with AI on data you can't share with it.

    Untraceable cloaks confidential values on your computer before any AI model sees the text — the model never receives the secret at all.