Metadata-Version: 2.4
Name: glyphmark
Version: 0.3.0
Summary: Post-inference invisible-Unicode text watermarking for AI-content provenance (EU AI Act Art. 50)
Project-URL: Homepage, https://textmark-65hh.onrender.com/
Author-email: Tanner Mathison <tanner.mathison@gmail.com>
License-Expression: Apache-2.0
License-File: LICENSE
Keywords: ai-generated,compliance,ed25519,eu-ai-act,hmac,provenance,steganography,unicode,watermark,zero-width
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Legal Industry
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Security :: Cryptography
Classifier: Topic :: Text Processing :: General
Classifier: Typing :: Typed
Requires-Python: >=3.10
Requires-Dist: cryptography>=42.0
Provides-Extra: dev
Requires-Dist: build; extra == 'dev'
Requires-Dist: matplotlib>=3.8; extra == 'dev'
Requires-Dist: pytest>=8.0; extra == 'dev'
Requires-Dist: twine; extra == 'dev'
Provides-Extra: web
Requires-Dist: flask>=3.0; extra == 'web'
Requires-Dist: gunicorn>=21.0; extra == 'web'
Requires-Dist: waitress>=3.0; extra == 'web'
Description-Content-Type: text/markdown

# glyphmark — post-inference text watermarking via hidden Unicode

`glyphmark` embeds an **invisible, keyed, tamper-evident signature** into text
*after* a model has produced it. The visible text is byte-for-byte unchanged,
but the result now carries:

1. **the request time** (minute-resolution, UTC), and
2. **a robust "this was AI-generated" proof** — an HMAC tag that only the holder
   of the secret key can mint, so it is detectable after the fact and hard to
   reproduce, and
3. **an optional literal message** (e.g. *"this has been text watermarked by
   Tanner"*).

It is designed to survive copy/paste, reformatting, splicing, rearrangement and
deletion, and to be detectable inside any ~100-word window.
This allows adding loose provenance signals to text when model inference is not available.

**Pros**
- Works with any length string.
- Durable under typical editing (but not against removal efforts).
- Extremely low probability of false positives.
- Can embed multiple marks showing provenance across different providers.

**Cons**
- Any stripping of special characters removes the signal (possibly a pro from a user-choice perspective, but a con from a surety perspective).
- Character count balloons to either accomplish (verbose messaging, resilient cryptography, etc. increase the character counts).
- Doesn't work for every text editor (e.g. special characters show up on Mac Core Text, etc.).
- Embedding a message does not itself prove the truth of the message.

---

## Install

```bash
pip install glyphmark
```

The only runtime dependency is `cryptography` (for Ed25519). HMAC mode is
stdlib-only.

## Use as a library

```python
from glyphmark import Watermarker
import os

# Configure once with YOUR secret key, then mark/verify many times.
wm = Watermarker(key=os.environ["GLYPHMARK_KEY"].encode())

marked = wm.embed(model_output, message="generated by Acme LLM").text   # send this on
result = wm.detect(some_text)            # -> DetectResult
if result.detected:
    print(result.timestamp_iso, result.message)
```

Public-key provenance — a third party can **verify but not forge** (the right
choice if you expose a detector):

```python
signer   = Watermarker.ed25519()         # signer holds the private key
marked   = signer.embed(text).text
pub      = signer.ed_public              # publish this (32 bytes)

verifier = Watermarker.verifier(pub)     # holds only the public key
verifier.detect(marked).detected         # True
```

One-shot functional API is also available: `from glyphmark import embed, detect`.

### Marking already-marked text (idempotency + multiple marks)
`embed()` is **not** silently idempotent — by default it won't double-mark.
`on_existing` controls the behavior when the text already carries a mark
verifiable under your key:

```python
glyphmark.is_marked(text, key=...)          # bool

wm.embed(text, on_existing="skip")     # default: if already marked by this key, no-op
wm.embed(text, on_existing="replace")  # strip ALL hidden chars, then mark afresh (one clean layer)
wm.embed(text, on_existing="add")      # deliberately layer another mark (multi-party / multi-stage)
```

Every mark gets a random 2-byte **mark id**, so any number of independent marks
can coexist and the detector reports each one separately:

```python
res = detect(text, key=...)
res.num_marks          # e.g. 3
for m in res.marks:    # each: m.timestamp_iso, m.message, m.core_copies
    ...
```

This is also why two marks' messages never get mixed together.

### Content binding (anti-transplant) — `bind_content=True`
A packet on its own proves *"minted with this key"* wherever it lands — so the
hidden characters copied out of a marked document and pasted into unrelated
text still validate, and a detector would appear to vouch for text the model
never produced. Content binding closes that gap:

```python
res = wm.embed(text, bind_content=True)          # or: glyphmark embed --bind-content
out = wm.detect(res.text)
out.changed_regions    # the spans whose text differs from what was marked
out.content_match      # fraction of REGIONS unchanged (None if not bound)
out.vouched_chars      # how much text the matching anchors actually cover
out.vouched_fraction   # ...as a share of the document. READ THIS ONE (see below)
out.content_match_min  # worst score across ALL marks -- prefer over content_match
out.binding_stripped   # True: a bound mark is present but its anchors are gone
```

Each CORE packet carries a 2-byte **anchor**: a keyed digest of the previous
`ANCHOR_CHARS` characters of normalized visible text, protected by the packet's
MAC/signature so it cannot be altered to match new surroundings. Each anchor
therefore reports on one **region**, and the detector tells you which regions
changed rather than issuing a verdict:

```
content bind : 20/22 regions unchanged since marking (~91%)
vouches for  : 1465 of 1631 visible chars (~90%) -- the rest is NOT covered by any anchor
changed      : chars 126-227, chars 205-306
```

Cosmetic edits are free — case, punctuation and reflow normalize away — while
changing the letters marks the regions covering them, and packets pasted into
unrelated text mark every region. Overhead is 2 bytes per CORE packet
(+16 hidden chars in word-safe, +2 in compact) **regardless of window width**:
a wider window digests more text into the same 2 bytes.

#### Sizing the window — the one real knob
`ANCHOR_CHARS` (default **100**) sets how much text each anchor covers. Override
it with `config.ANCHOR_CHARS`, `--bind-chars N`, `Watermarker(bind_chars=N)`, or
the `bind_chars=` argument of the functional `embed()` / `detect()`. Set it relative to the gap between
packets — roughly `word_interval × 6` characters, so ~60 at `density="high"`,
~110 at `"medium"`, ~180 at `"low"`. A window at least that wide makes the
anchors *tile* the document. Measured on English prose at the default density:

| Window | reword 10% | delete 20% | coverage | forger must borrow |
|---|---|---|---|---|
| 40 chars | 0.51 | 0.79 | 55% | ~23% |
| **100 chars** | **0.19** | **0.49** | **~100%** | **~44%** |
| 200 chars | 0.04 | 0.21 | ~100% | ~71% |

Wider covers more of the document and costs a forger more genuine text;
narrower survives heavier editing and localizes changes more precisely. Both
sides must use the same value, or every region reports as changed.

Note what the wide default does to the number: at 100 characters a region spans
~17 words, so rewording 10% of a document marks ~80% of regions — and
`0.9¹⁷ ≈ 0.17` says that is *correct*, not a false alarm. Read `content_match`
as "how much of the marked text is unchanged, at region granularity", never as
a trust score.

Two properties stop an attacker from *searching* for text that satisfies a
stolen packet's anchor:

* **The anchor is keyed** (HMAC under a key derived from your watermark key),
  so nobody can compute what a candidate wording would hash to and search for
  filler that fits. In **Ed25519 mode the anchor key comes from the public
  key** — unavoidable, since a public verifier must recompute it — so there the
  search *is* available offline, which is why Ed25519 anchors are **8 bytes**
  rather than 2. (At 2 bytes a collision search found matching filler in 77,003
  tries / 0.29 s; at 8 bytes it is ~2⁶⁴.)
* **It binds whole words**, not a sketch of them, so a satisfying phrase is the
  real one rather than any wording with the same initials.

#### What binding does *not* stop — read before relying on the result
Binding is **local**, and no window width changes that. An anchor attests to
`ANCHOR_CHARS` characters; splicing a packet together with those characters
into a new document is indistinguishable from *quoting* a marked passage,
because it is the same operation. A forger who cuts
`<the bound text><the packet>` fragments out of a marked document and scatters
them through their own prose gets those anchors to match. No key is needed —
bound packets are identifiable by length alone.

What the window buys is a **cost**, not a proof. Measured at the default 100
characters, a forger carrying too little gets nothing, and matching every
anchor takes ~44% of the forgery being genuine borrowed text:

| Carried per packet | Regions matching | Forgery that is real text |
|---|---|---|
| 25 chars | 0 / 6 | 11% |
| 50 chars | 0 / 6 | 20% |
| 150 chars | 5 / 6 | 44% |
| 250 chars | 5 / 6 | 62% |

Raising `ANCHOR_CHARS` raises that fraction and lowers edit-robustness by
exactly the same mechanism — they are the same question asked twice.

This is also why coverage is reported separately. A document assembled from a
few genuine fragments can show every anchor matching while covering almost
none of itself, so check `vouched_fraction` before reading `content_match` as
a statement about the document.

Notes:
* **Off by default.** With the flag off, embedding is byte-identical to
  v0.2.x (pinned by `tests/test_compat_v020.py` against the published wheel).
  A **bound** mark needs a v0.3.0+ detector — both its CORE and MESSAGE
  packets use new MAC domains, so a v0.2.x detector reports bound text as
  unmarked. Upgrade detectors before enabling the flag on the embed side.
* Multiple marks: `content_match` describes the *primary* mark only, so a
  genuine mark of your own could otherwise mask someone else's transplanted
  one — prefer `content_match_min`. That is not airtight either: two marks that
  cannot be told apart get **merged into one** and report a single blended
  ratio. This happens when Ed25519 marks (which carry no mark id) share a
  timestamp *and* message — an attacker can read both from a published document
  with only the public key — and, at odds of 1 in 65,536, on an HMAC mark-id
  collision.
* **There is deliberately no "suspicious below X%" threshold.** Any cutoff
  would have to separate heavy editing from transplantation, and the anchors
  do not carry that information — a genuine document reworded enough lands
  wherever a forgery lands. The detector reports which regions changed and how
  much text the matching anchors cover, and leaves the judgement to you. The
  one unambiguous case, **no** region matching at all, is called out on its own
  because it needs no tuning.
* MESSAGE packets carry no anchor (it would double message overhead) but are
  emitted under their own MAC domain when the mark is bound, so deleting the
  bound CORE packets is reported as `binding_stripped` rather than passing the
  mark off as never-bound.
* To customize *what* the mark is bound to (window size, digest width, a
  looser signature), edit the clearly marked block in
  [`glyphmark/anchor.py`](glyphmark/anchor.py) — embedder and detector share
  those functions, so they stay in sync.

### Key management (read this for compliance use)
Provenance security rests entirely on key secrecy.

* **HMAC mode** — pass your own `key=` (or set `GLYPHMARK_KEY`). If you pass
  nothing, the library falls back to a **public demo key and emits a warning**;
  watermarks made with it are forgeable. Never use it in production.
* **Ed25519 mode** — keep the private seed secret; distribute only `ed_public`.
  Anyone can verify with the public key; nobody can forge without the private
  one. This cleanly separates "anyone can check" from "only you can mint."

### CLI

```bash
echo "The model wrote this." | glyphmark embed --message "by Acme" > out.txt
glyphmark detect -i out.txt
# without installing: python -m glyphmark embed ...
```

`detect` exits **0** = watermark found, **1** = none found, **2** = authentic
packets of which **not one** region matches the text they sit in — or a bound
mark whose anchors are all gone (`binding_stripped`, which heavy deletion also
causes). Partial change is reported in the summary rather than in the exit
code, since no cutoff can separate heavy editing from a partial transplant. The
verdict is judged on the *primary* mark only — anyone can append invisible
characters lifted from another document, and that must not let a stranger
control your exit code.

### The four fields

| Watermarker | |
|---|---|
| **Input text** | the text to watermark |
| **Output text** | the watermarked text (visible chars unchanged) |
| **Watermark message** *(optional)* | literal string embedded in the hidden layer |

| Detector | |
|---|---|
| **Input text** | watermarked (possibly edited) text |
| **Output text** | the recovered signal: AI-generated proof + request time (+ message) |

---

## How it works

Two independent layers:

### 1. Robust keyed packets (the real signal)
Payload bytes are encoded as invisible characters. There are two encodings:

* **word-safe (default)** — binary over two zero-width characters,
  ZWNJ (`U+200C`) = 0 and ZWJ (`U+200D`) = 1, one bit each (8 chars/byte), with
  WORD JOINER (`U+2060`) as decoy noise. Invisible in **Microsoft Word, Google
  Docs**, browsers and chat, and survives a Word copy/paste round-trip. We
  *avoid* `U+200B` (ZWSP) because **Word silently strips it on paste**, and
  variation selectors because they render as visible boxes (tofu) in Word/Docs.
* **compact** (`--compact` / site toggle) — one **variation selector** per byte
  (`U+FE00–FE0F`, `U+E0100–E01EF`); ~4× denser, but **shows as boxes in
  Word/Docs**. Use only when the text will live in browsers/chat.

Packets are short (**10 bytes**) and self-contained, so a single surviving copy
recovers everything. Each carries a 2-byte **mark id** so a mark's packets group
together (enabling any number of coexisting marks):

```
CORE        mark_id(2) || ts(3) || HMAC(key,"C"||mark_id||ts)[:5]              provenance + time
MESSAGE     mark_id(2) || idx(1) || total(1) || msg(2) || HMAC(key,"M"||…)[:4]  one message chunk
CORE-BOUND  mark_id(2) || ts(3) || anchor(2) || HMAC(key,"B"||…)[:5]            + content binding
```

* The **HMAC tag is the provenance proof**: forging a packet that validates is
  ~`2**-40` per attempt without the key. Verifying it after the fact only needs
  the key. → *detectable, hard to reproduce.*
* Packets are **repeated throughout** the text (a CORE packet roughly every
  10–30 words depending on density), so any ~100-word window contains several.
* Each packet is **XOR-whitened** with a key-derived keystream, hiding any
  constant header/structure from someone inspecting the raw bytes.
* The detector **slides a window** over the invisible stream (10 bytes in
  compact, 80 zero-width bits in word-safe), keeps the windows whose MAC
  validates, and **groups them by mark id** — so it recovers every mark
  regardless of order, and a single intact packet suffices. This is why
  splicing/rearranging and block deletion barely dent it. (Word-safe packets are
  longer, so uniform per-character deletion hurts them more than compact — block
  edits are unaffected.)

### 2. Context-driven obfuscation (the confusion)
A large rule engine (`rules.py`) reads features of the
surrounding **visible** text — word-length parity, vowel/consonant balance,
digits, capitalization, punctuation, position mod *N*, keyed hashes of each word,
sentence index, and ~25 interacting rules — to decide **where** packets land
(keyed jitter), **how many decoy characters** to scatter, and **which decoy
alphabet** to draw from (variation selectors vs. a family of zero-width
characters). The decoys are key-derived, random-looking **garbage** that never
carries a valid MAC, so:

* an outside observer sees a complex, context-dependent mess that is hard to
  reverse-engineer or reproduce, while
* the legitimate detector simply ignores everything that does not validate.

Crucially, the obfuscation layer **never touches the bytes of a real packet**,
so it adds confusion without costing robustness. This is mostly for fun.

---

## Robustness

* **Block deletion / splicing / rearrangement** (deleting whole words,
  sentences, paragraphs; reordering) — the realistic editing case — recovers
  essentially **100%** even when most of the document is removed, because
  surviving packets stay intact and order doesn't matter.
* **Uniform per-character deletion** (a harsh, less realistic channel) is where
  the two encodings diverge sharply. In `--compact` a packet is 10 characters,
  so recovery holds at ~100% to 20% deletion. In the default word-safe encoding
  a packet is 80 zero-width characters and `0.9**80` is effectively zero, so it
  collapses between 5% and 10%. Block edits — the realistic case — are
  unaffected either way.
* **Wrong key / no key** → nothing validates (the provenance guarantee).

* **Packet transplantation** (copying the hidden characters into other text) —
  undetectable by design without binding; with `bind_content=True`, packets
  copied *alone* mark every region changed while provenance still holds.
  Packets copied *with* their bound text are not — see the limits above.

Run `python make_figure.py` to regenerate the plot, and `pytest tests/` for the
round-trip, robustness, binding and v0.2.x-compatibility suite (68 tests).

### Honest limitations
* Survival to copy/paste/reformatting depends on the destination preserving
  invisible code points. Most chat apps, docs and clipboards do; some
  sanitizers (e.g. systems that strip non-printing characters, or re-typing the
  text) will remove them — no invisible scheme can survive that.
* Provenance security rests entirely on **key secrecy** — keep `GLYPHMARK_KEY`
  server-side. The obfuscation layer is defense-in-depth, not a substitute.
* Content binding is *local and graded*, not a signature over the document: it
  answers "were these packets minted over the text they now sit in?", not
  "is this the whole original document". A document assembled from genuine
  fragments still matches, and in Ed25519 mode the anchor key is public. Treat
  the result as evidence about the `vouched_chars` it covers, never as proof
  about the document.
* Overhead is invisible but real: a few hundred invisible chars for a paragraph.
  Provenance-only is cheapest; embedding a literal message multiplies it
  (each 2-byte chunk is repeated). Use `--density low` / `--no-decoys` to minimize.

---

## Configuration
* `GLYPHMARK_KEY` — secret HMAC key (env var). Use the same key to embed/detect.
* `--density {low,medium,high}` — redundancy vs. character count trade-off.
* `--no-decoys` — disable the garbage/obfuscation characters.
* `--bind-content` — bind packets to the surrounding words (anti-transplant;
  see above). Customize the binding in `glyphmark/anchor.py`.

---

## Try it
**Live demo:** https://textmark-65hh.onrender.com/ — paste text, watermark it,
edit it, and detect it in your browser.
