richeddy shared this post ยท 2h ago
Jonas Geiping

Why am I being baited by watermark misinformation on this app, is it 2023 again?

A small FAQ:

  1. ๐—ช๐—ต๐—ฎ๐˜'๐˜€ ๐—ฎ ๐˜๐—ฒ๐˜…๐˜ ๐˜„๐—ฎ๐˜๐—ฒ๐—ฟ๐—บ๐—ฎ๐—ฟ๐—ธ? -- A modification of the LLM sampling algorithm that, if there are multiple ways to write something, will pick one that agrees with a pseudorandom key. This is a local, invisible signature hidden in the way phrases are used in any LLM text that persists when text is copied.

  2. ๐——๐—ผ๐—ฒ๐˜€ ๐˜๐—ต๐—ถ๐˜€ ๐—บ๐—ฎ๐—ธ๐—ฒ ๐˜๐—ต๐—ฒ ๐˜๐—ฒ๐˜…๐˜ ๐˜„๐—ผ๐—ฟ๐˜€๐—ฒ? -- A good implementation is 'undetectable' (in polynomial time), meaning: If you do not have the private key, then neither you, the model itself, or pangram could detect that this is happening.

  3. ๐—ช๐—ถ๐—น๐—น ๐˜๐—ต๐—ถ๐˜€ ๐—ฏ๐—ฟ๐—ฒ๐—ฎ๐—ธ ๐˜๐—ต๐—ฒ ๐—บ๐—ผ๐—ฑ๐—ฒ๐—น'๐˜€ ๐—ฟ๐—ฒ๐—ฎ๐˜€๐—ผ๐—ป๐—ถ๐—ป๐—ด? -- Because Ant already encrypts the model's reasoning, they can just not watermark the model's internal reasoning, leaving the thinking unaffected.

  4. ๐—ช๐—ถ๐—น๐—น ๐˜๐—ต๐—ถ๐˜€ ๐—บ๐—ฎ๐—ธ๐—ฒ ๐˜๐—ต๐—ฒ ๐—บ๐—ผ๐—ฑ๐—ฒ๐—น ๐—น๐—ฒ๐˜€๐˜€ ๐—ฐ๐—ฟ๐—ฒ๐—ฎ๐˜๐—ถ๐˜ƒ๐—ฒ/ ๐—บ๐—ผ๐—ฟ๐—ฒ ๐˜€๐—ฎ๐—บ๐—ฒ-๐˜†? -- If anything this (marginally) increases entropy across different generations, so it will make model outputs slightly more varied.

  5. ๐—•๐˜‚๐˜ ๐—œ ๐—ฐ๐—ฎ๐—ป ๐—ท๐˜‚๐˜€๐˜ ๐—ฟ๐—ฒ๐—บ๐—ผ๐˜ƒ๐—ฒ ๐—ถ๐˜ ๐—ฝ๐—ฎ๐—ฟ๐—ฎ๐—ฝ๐—ต๐—ฟ๐—ฎ๐˜€๐—ถ๐—ป๐—ด? -- Absolutely! But, judging from the amount of writing on the web that already unmistakably sounds like Claude, most people likely will not bother.

5b: Also, not any paraphrase will work. To remove (for example) a k=5-minhash watermark completely from a long document, you need to make sure none of the original 2-grams, 3-grams, 4-grams, 5-grams and 6-grams of the text remain.

  1. ๐—ช๐—ถ๐—น๐—น ๐˜†๐—ผ๐˜‚ ๐—ถ๐—ป๐—ฎ๐—ฑ๐˜ƒ๐—ฒ๐—ฟ๐˜๐—ฒ๐—ป๐˜๐—น๐˜† ๐—ฐ๐—ผ๐—ฝ๐˜† ๐˜๐—ต๐—ฒ ๐˜„๐—ฎ๐˜๐—ฒ๐—ฟ๐—บ๐—ฎ๐—ฟ๐—ธ? -- No, with a good implementation the space of possible realizations of the key is too large to memorize.

  2. ๐—ช๐—ถ๐—น๐—น ๐˜๐—ต๐—ถ๐˜€ ๐—ฎ๐—น๐—น๐—ผ๐˜„ ๐—–๐—น๐—ฎ๐˜‚๐—ฑ๐—ฒ๐˜€ ๐˜๐—ผ ๐—ถ๐—ฑ๐—ฒ๐—ป๐˜๐—ถ๐—ณ๐˜† ๐—ผ๐˜๐—ต๐—ฒ๐—ฟ ๐—ถ๐—ป๐˜€๐˜๐—ฎ๐—ป๐—ฐ๐—ฒ๐˜€ ๐—ถ๐—ป ๐—ฎ ๐˜€๐˜„๐—ฎ๐—ฟ๐—บ? -- The watermark will 'appear' like random sampler fluctuation to the model and would not be detectable. But, if an agent gets hold of a detector endpoint, it can absolutely use the watermark to ID other Claude agents (not that it would have trouble noticing them based on their writing as of today).

  3. ๐—ช๐—ถ๐—น๐—น ๐˜๐—ต๐—ถ๐˜€ ๐—ฑ๐—ฒ๐˜๐—ฒ๐—ฐ๐˜ ๐—ฑ๐—ถ๐˜€๐˜๐—ถ๐—น๐—น๐—ฎ๐˜๐—ถ๐—ผ๐—ป? -- By default, no. If the watermark is set up to be 'undetectable' (as assumed above), it will not be picked up in training by other models. For that to happen, the watermark needs to be detectable by ML algorithms.

  4. ๐—ช๐—ถ๐—น๐—น ๐˜๐—ต๐—ถ๐˜€ ๐—บ๐—ฎ๐—ธ๐—ฒ ๐—ฃ๐—ฎ๐—ป๐—ด๐—ฟ๐—ฎ๐—บ'๐˜€ ๐—ท๐—ผ๐—ฏ ๐—ฒ๐—ฎ๐˜€๐—ถ๐—ฒ๐—ฟ? -- By default no, this is a separate avenue to detection. But, they might collaborate with Anthropic which would allow them to detect the watermark as well and show a watermark score next to their text detection score.

  5. Bonus: All aside, is this a good idea? I don't know. The companies are doing it to follow the writing of the EU AI act, which was written based on 2024 information and when the field looked very different, and threat models were focused much more on slop/propaganda (like the Kokotajlo 2026 prediction). The actual 2026 looks quite a bit different.

427