FIU’s Mustafa Ocal says the Claude watermark biases word choice toward hidden patterns only Anthropic’s scanner can read. Older models lack it. #Watermark #Claude #FIU #AIDetection
The Claude watermark hides in plain sight, and a Florida International University expert says that is by design.
An FIU explainer, distributed through Community Newspapers, breaks down how Anthropic’s watermark works. It also explains why no reader can find it by eye.
How Does the Claude Watermark Work?
Start with how language models write, says Mustafa Ocal, a professor at FIU’s Knight Foundation School of Computing & Information Sciences who researches AI detection. “AI models don’t ‘think.’ They predict.”
At each step, the model lists words that could come next and attaches a probability to each. Then it picks one at random, weighted by those odds.
Consider a one-sentence story about fishing. The model might write that the rod bent “sharply” as the bass fought the line. “Sharply” was not the only word that worked. It was drawn from a pool of highly probable options.
Ocal compares this to drawing marbles from a bag. The watermark changes the bag. “Claude’s watermark works by reducing randomness in word choice,” the explainer says.
What Does a Watermarked Sentence Look Like?
Ocal uses a hypothetical. Ask Claude for the greatest food in the world, and it might write that pizza is one of the most “satisfying” foods there is.
Before the watermark, the odds for that slot might read like this: delicious 21%, fulfilling 19%, satisfying 17%, gratifying 14%. With the watermark, satisfying jumps to 27% and gratifying to 23%. Delicious falls to 8%.
Ocal describes it as dyeing some words blue and others red. Blue words appear more often. Red words appear less often.
In a longer sample, a high ratio of blue to red points to Claude. In a novel example Ocal examined, 80% of the dyed words were blue — an “early indicator” that Claude may have written it, though he says a larger sample would be needed to confirm.
Why Can’t Readers Spot It?
The blue and red sets change with context. “We don’t know which words to look for,” Ocal says.
He likens it to drawing blackjack half the time at a casino. A human is extremely unlikely to land on blue words at that rate, because so many other word choices are available.
Each blue word looks normal, because each has a close synonym a human might have used. Only Anthropic’s scanner sees the pattern, which Ocal says looks very visible to a machine that can see blue and red.
Commercial AI detectors work differently. They are trained to separate AI writing from human writing by style. The watermark concerns only word choice.
Light proofreading is unlikely to leave a trace, Ocal says. If Claude changes word choice, a trace becomes possible.
Who Benefits From a Watermark?
Ocal calls watermarking “an excellent idea.” Two uses stand out in the explainer.
The first is hiring. “They’ll have to disclose when they use Claude to make hiring decisions,” Ocal says of employers — his own view of where disclosure norms are headed, not a stated legal requirement. He argues that is good for applicants.
The second is evidence. Ocal says that if the U.S. Department of Justice had this scanner, it could check whether a document was generated by Claude.
Both uses depend on a detection tool. MAIN reported in August that Anthropic plans a watermark-detection API but has not released details.
What Are the Limits?
The watermark applies only to Claude models launched on or after August 2, 2026, according to MAIN’s earlier report on Claude’s invisible watermark. Older models, which most users still run, do not carry it.
That report also noted that detection is weaker for short or highly factual passages, and that a complete rewrite can remove the pattern.
FIU, a Miami university of about 55,000 students, has built a research profile in AI security. Its CIERTA center studies adversarial machine learning and LLM security, as MAIN has covered.
The watermark is a probability signal, not proof of authorship. Executives who rely on it should treat it as one input among several.
Sources: