The deepfake that was, and the one that wasn’t
In August 2006, a Reuters photographer named Adnan Hajj filed a picture from Beirut: thick black smoke rising over the city after an Israeli airstrike. Within hours of its publication, American blogger Charles Johnson grew suspicious. The smoke looked wrong to him, as if someone had cloned it in a photo editor. On Johnson’s tip, Reuters investigated, then pulled more than nine hundred of Hajj’s photographs, fired him, and let his editor go as well. The agency tightened its digital photography standards after publishing its complete findings in January 2007: now it allows only cropping and exposure adjustments, the same things a photographer could do in a darkroom.
In March 2021, police in Pennsylvania arrested Raffaela Spone for allegedly making deepfake videos of her daughter’s cheerleading rivals. She had, the story went, fabricated photos and footage of the girls drinking, vaping and in states of undress. The detective on the case admitted that he had spotted the fakery "by eye." The story swept through the national press overnight, but prosecutors dropped the deepfake claims by May after the only video that was made public turned out to be real.
"Her reputation right now is less than mud. They have ruined her life, there’s no question about it. … On Twitter, they have already convicted her. She’s always going to be labeled the 'deepfake mom' or … a 'criminal mastermind.'"
— Robert J. Birch, attorney for Raffaela Spone, quoted in: Drew Harwell, "Remember the 'deepfake cheerleader mom'? Prosecutors now admit they can’t prove fake-video claims,", The Washington Post, May 14, 2021
The response to a suspected deepfake was equally confident in both these stories, even though a forgery only existed in one. A blogger might know more about photo editing than a detective, and Reuters investigated the circumstances before taking action, but human judgment prompted each case.
How would each incident have gone with a tool that could reliably spot AI traces in photos or videos?
That tool now exists: on August 14, 2026, Anthropic announced that all text Claude processes will carry an invisible watermark. We need to understand what that change does and who it serves.

You now have a watermark you can’t read…
A watermark in a text isn’t an author’s signature. Anthropic is blunt about this:
"A watermark only helps test whether Claude might have produced or processed the content. It doesn’t say anything about ownership or authorship, and doesn’t change a user’s rights under our terms."
— Anthropic, How Claude’s text watermark works, August 14, 2026
The mechanism rests on SynthID, a method developed at Google DeepMind. When the model generates text, it doesn’t pick each next token at random. It follows a pattern set by a hidden key, and the detector later checks whether the sequence of choices matches those instructions. Anthropic offers an analogy: picture a game of Monopoly in which the digits of pi stand in for the dice. The result looks like chance, but it can be reproduced and verified.
Anthropic itself lists what SynthID can’t catch:
- Short texts, where the signal is too faint to show up.
- Factual passages where the next word is the only one possible, like "Isaac Newton’s most famous work was called Principia Mathematica."
- Edits to someone else’s prose. If Claude corrects only the grammar, the watermark has nowhere to hide but in the corrections, and there may be too few of them.
Finally, the limitation that matters most:
"It cannot distinguish 'Claude wrote this' from 'Claude heavily edited this.'"
— Anthropic, How Claude’s text watermark works, August 14, 2026
None of these statements mean Anthropic’s dodging responsibility. The company’s representatives give an honest account of how SynthID works and what it can actually do.
An editor needs to know where the line falls: did the author hand the machine only a proofreading pass, or did Claude write the entire draft, first letter to last? A journalist has to establish where a document came from. A scholar wants to shield a paper from the charge of automated authorship. A student wants to prove that the essay is their own.
The watermark gives them none of this. It can’t tell structural generation from a light polish. It doesn’t vouch for your authorship, and it does nothing for your work. It merely records that an algorithm once touched the text. Even that trace is off-limits to you, since you almost certainly don’t have a scanner.

…because such marks exist to make a prosecutor’s life easier, not yours
The history of labeling doesn’t begin with algorithms.
In the mid-1980s, Xerox built a hidden feature into its color laser printers: every page came out stamped with a constellation of tiny yellow dots, invisible to the naked eye, officially to deter counterfeiters. The dots encoded the printer’s serial number and the date and time of printing. In October 2005, the Electronic Frontier Foundation exposed the system and published a tool for decoding it.
In May 2017, Reality Winner, a 25-year-old NSA contractor in Augusta, Georgia, printed a five-page report on Russian interference in the American election and sent it to The Intercept. Before publication, the editors handed a scan of the document to the government to authenticate it, not knowing that the dots had survived the scan. The dots recorded a model 54 printer with serial number 29535218, and a moment: May 9, 2017, 6:20.
"The document leaked by the Intercept was from a printer with model number 54, serial number 29535218. The document was printed on May 9, 2017 at 6:20. The NSA almost certainly has a record of who used the printer at that time."
— Robert Graham, How The Intercept Outed Reality Winner, Errata Security, June 5, 2017
Winner was arrested on June 3, two days before The Intercept published. A tool built to fight counterfeiting had been turned on a whistleblower.
This is a structural feature of any labeling system: it always serves whoever controls the verification. The stated purpose and the actual use rarely coincide, because the labeling tool ends up in the same hands as the detector.

Yet this protection collapses at a simple paraphrase
The tool is technically fragile, too. Google DeepMind open-sourced the SynthID code in October 2024. In August 2026, Anthropic rolled out its watermark with no regional restrictions at all.
How durable is the trace? A study called Watermark under Fire (Liang et al., accepted at EMNLP 2025) put SynthID through its paces. On untouched text, detection accuracy is 99.8%. Copy a passage straight out of the Claude window, and the detector will catch you almost every time. Under moderate paraphrasing, however, accuracy drops to 49.8%, no better than a coin toss. This is no fluke; it’s a general property of the method.
Meanwhile, open-weight models such as Llama and DeepSeek, which for architectural reasons carry no watermark, can run locally. If even one major player stays outside the system, the structure means this model becomes a way around the entire control apparatus, even if the creators never intended that.

Why AI providers were in no hurry to press the button
Technical fragility is only half the problem. The other half is a deliberate corporate choice. At OpenAI, a finished watermarking tool has been sitting in a drawer since at least the summer of 2023.
"It’s just a matter of pressing a button."
— one of the people familiar with the matter, quoted in: Deepa Seetharaman & Matt Barnum, "There’s a Tool to Catch Students Cheating With ChatGPT. OpenAI Hasn’t Released It.", The Wall Street Journal, August 4, 2024
The button was never pressed. In April 2023, OpenAI ran two surveys. The first, among the general public, found people favoring watermarks by four to one. The second, among ChatGPT’s own active users, told a different story: 69% feared false accusations, and 30% said they would switch to a competitor that didn’t use watermarks. Internal company documents obtained by The Wall Street Journal record where that contradiction left the company:
"Our ability to defend our lack of text watermarking is weak now that we know it doesn’t degrade outputs."
— internal OpenAI documents, quoted in: Deepa Seetharaman & Matt Barnum, "There’s a Tool to Catch Students Cheating With ChatGPT. OpenAI Hasn’t Released It.", The Wall Street Journal, August 4, 2024
As of August 2026, OpenAI’s button remained unpressed. Anthropic pressed its own and rolled out watermarking worldwide, but the detector is available only to regulators and a narrow circle of organizations. The public’s appetite for catching AI red-handed still hasn’t gone anywhere, however, and crude knockoffs have moved quickly to fill the vacuum the industry left.

Leaving you alone with crude substitutes that punish bad English
Turnitin and GPTZero don’t read watermarks. It’s worth being precise here: they don’t check whether a text carries a SynthID signal or any other mark. They analyze the statistical patterns of a text and compare them with how people write—or, to be exact, with how people usually write when English is their first language.
In 2023, researchers at Stanford tested seven AI detectors on two sets of human-written texts: essays by American eighth-graders and essays from the TOEFL exam, which tests non-native English speakers:
"These detectors consistently misclassify non-native English writing samples as AI-generated, whereas native writing samples are accurately identified."
— Weixin Liang et al., "GPT detectors are biased against non-native English writers", Patterns, 2023
In 61.22% of cases, the detectors branded TOEFL essays the work of an algorithm, while nearly always correctly identifying the American schoolchildren’s essays as human. At least one detector perceived AI in 97.8% of TOEFL texts.
The detectors stumble even on their own terms: in another study, Debora Weber-Wulff’s group tested fourteen detectors and didn’t get more than 80% accuracy from any of them. Running the AI-generated text through an automated paraphrasing tool dropped average accuracy to 26%. (For a closer look at how detectors work and who pays for their false alarms, see "Are AI Detectors Accurate? How They Get It Wrong and Who Pays.")
The Stanford authors spelled out the consequence:
"Paradoxically, to evade false detection as AI-generated content, these writers may need to rely on AI tools to refine their vocabulary and linguistic diversity."
— Weixin Liang et al., "GPT detectors are biased against non-native English writers", Patterns, 2023
That is the researchers’ conclusion, not a recommendation.
In 2023, a professor at Texas A&M refused to grade his graduating seniors’ final papers. To check them, he fed the papers to ChatGPT and asked: "Did you write this?" ChatGPT answered "Yes" without checking any watermark. The professor later reconsidered his decision.

The pros get the real tool
Anthropic’s Detection API isn’t a public feature. A select group of professionals gets it at launch: regulators, law enforcement, newsrooms, fact-checkers, independent researchers and educational institutions in the EU. Teachers with a Turnitin account can’t use it, and neither can online readers.
When a studio sends a new film to reviewers before its premiere, it adds a unique mark to each copy that only those in the company can see, so it can trace any leaks back to their source. Anthropic’s Detection API works on exactly the same principle: it lets those with access and authority examine a text after the fact.
This tool’s existence isn’t a problem. The trouble begins when people think public AI detectors like Turnitin and GPTZero can do the same job.
In June 2023, The Washington Post reported on teachers across the country who were using AI detectors to check student work. Soheil Feizi of the University of Maryland, one of the leading researchers in the field, argues that a detector used on students should produce no more than 0.01% false positives. Is that figure within reach?
"At this point, it’s impossible. And as we have improvements in large-language models, it will get even more difficult to get even close to that threshold."
— Soheil Feizi, University of Maryland, quoted in: Geoffrey A. Fowler, "Detecting AI may be impossible. That’s a big problem for teachers.", The Washington Post, June 2, 2023
Feizi was speaking of detection without watermarks. Confusing public AI detectors with the Detection API and similar tools is costly, and the people who mix them up aren’t the ones who’ll pay.

The rules will appear after you’ve paid
Reuters didn’t invent a special algorithm for catching doctored photographs. It simply sat its editors down with its photographers and agreed on the line between acceptable correction and falsification. The process is slow and awkward, and it comes with no guarantee, but it has held to this day.
A watermark doesn’t draw that line. It only records that whoever created a text used an algorithm for at least part of it, and only shows that to those who have the detector, the access and the authority. For a regulator, it’s a useful tool. For an editor at a publishing house or a schoolteacher, it doesn’t exist.
Raffaela Spone walked out of court with three years of probation, convicted not of deepfakes but of harassment using anonymous phone numbers. By that time, the headlines had already circled the globe. The detective who had spotted the supposed forgery "by eye" never explained what his certainty rested on.
Norms are always rewritten after someone has paid the price. Reuters wrote its own in 2007, in the wake of Hajj. The AI industry is writing its own now, after it became clear that no one could do without them, and before it becomes clear who exactly they’re for.
Sources
References cited in this piece. Last verified on the published or revision date.
- 01
- 02
- 03
- 04
- 05
- 06
- 07
- 08
- 09
- 10