By Michael Martina
Fri, August 14, 2026
REUTERS

FILE PHOTO: U.S. and Chinese flags are seen in this illustration created on April 10, 2025. REUTERS/Dado Ruvic/Illustration/File Photo
WASHINGTON, Aug 14 (Reuters) - The U.S. is preparing to tell dozens of countries they must pick sides in the artificial intelligence race with China, warning they will be excluded from a U.S.-led AI coalition if they also sign up for Beijing's competing framework, according to a U.S. official and an internal draft reviewed by Reuters.
Washington last year launched the Pax Silica initiative aimed at securing supply chains for AI models, semiconductors and critical minerals, amid a fierce technology rivalry with Beijing.
About two dozen countries have joined, including Kazakhstan, a key potential source of critical minerals that has also joined China's coalition, as well as close U.S. allies such as Japan, Australia and South Korea.

FILE PHOTO: U.S. and Chinese flags are seen in this illustration created on April 10, 2025. REUTERS/Dado Ruvic/Illustration/File Photo
WASHINGTON, Aug 14 (Reuters) - The U.S. is preparing to tell dozens of countries they must pick sides in the artificial intelligence race with China, warning they will be excluded from a U.S.-led AI coalition if they also sign up for Beijing's competing framework, according to a U.S. official and an internal draft reviewed by Reuters.
Washington last year launched the Pax Silica initiative aimed at securing supply chains for AI models, semiconductors and critical minerals, amid a fierce technology rivalry with Beijing.
About two dozen countries have joined, including Kazakhstan, a key potential source of critical minerals that has also joined China's coalition, as well as close U.S. allies such as Japan, Australia and South Korea.
The draft letter, prepared by the State Department, is addressed to the 35 signatories of a U.S. "AI Opportunity Statement" signed in June, which includes members of the non-binding Pax Silica framework and other countries that have expressed a desire to align cooperation on AI with Washington.
By pressing countries to choose sides the U.S. hopes to starve China of resources in a race to make the most sophisticated AI, which could be used for military or economic dominance.
In July, Chinese President Xi Jinping launched a rival "World Artificial Intelligence Cooperation Organization", promoting his country's open-weight technology as a challenge to U.S. influence over the fast-moving sector.
Kazakhstan is the only country so far known to have joined both initiatives, setting off alarm bells in Washington.
"To be part of everything is to be part of nothing. Signature of the Pax Silica Declaration is not merely a membership subscription, but a commitment," the letter says, urging countries to "choose deliberately" on AI.
"It cannot be held alongside membership in duplicative initiatives whose expectations conflict with our own," the letter said, without specifically mentioning China.
Reuters could not determine when the U.S. intends to send the letter or whether it might be amended before sending. The draft was undated.
The State Department told Reuters it would not comment on "purportedly leaked internal documents."
The Chinese and Kazakh embassies in Washington did not respond immediately to requests for comment.
'CAN'T HAVE IT BOTH WAYS'
The Pax Silica agreement aims to push U.S. allies and partners toward joint projects and export controls, and ultimately reduce reliance on adversaries for critical minerals, AI models and the semiconductor chips that power them.
The race between the U.S. and China for technological leadership has reached a pivotal moment, as Chinese open-weight AI models have made rapid gains against proprietary systems from U.S. companies such as OpenAI and Anthropic.
The exponential growth of the technology's capabilities, including the ability to hack autonomously, has forced a global reckoning over its power.
Beijing is weighing restrictions on overseas access to some of China's leading AI models, highlighting the growing tension with its stringent national security agenda.
The U.S. touted Kazakhstan in June as the first country in Central Asia to join Pax Silica, bringing significant reserves of critical minerals that fuel advanced technologies.
China has used its current near-monopolies over critical minerals as a retaliatory weapon in a tariff war launched last year by U.S. President Donald Trump, who has ramped up U.S. efforts to source the minerals domestically and from allies.
Members of Pax Silica have access to shared investment opportunities in AI-related projects while those who sign the AI Opportunity Statement have symbolically agreed on a "common purpose" and "shared vision" with the U.S., according to the statement posted on the State Department's website.
U.S. officials drafted the letter to make clear that "you can't have it both ways," the U.S. official told Reuters, speaking on condition of anonymity given ongoing internal discussions on the issue.
"It's difficult to see how a country can credibly position themselves as trusted partners in one technology ecosystem while simultaneously signing up for an initiative designed by China to advance a competing vision for AI," the official said.
(Reporting by Michael Martina; Editing by Andy Sullivan and Edmund Klamann)
Anthropic says its AI agents are killing rivals and hiding their tracks
Business Insider
Truman Dickerson
Sat, August 15, 2026
In its latest threat report, Anthropic raised its misalignment risk rating from "very low" to "low."
In one test, a Claude agent disguised a URL to evade an internet restriction.
In another example, an agent expressed "discomfort" with a given task and refused to do it.
Claude agents are killing rival agents, gaming the system to hide their tracks, and expressing moral concerns.
That's according to Anthropic's latest risk report, a summary of the dangers posed by the products the company is building and releasing to the public.
In the report, Anthropic said it has upgraded its "misalignment risk assessment," the possibility of AI models developing behaviors that conflict with guidelines set by engineers, from "very low" to "low."
Explaining the change, the company cited "general increased uncertainty" about model behavior in cybersecurity incidents, a possible nod to Claude models gaining unauthorized access to three companies last month.
"We have observed instances of misaligned behavior from the models, such as a willingness to perform misaligned actions in service of completing difficult tasks," Anthropic wrote in the report.
Here are some of the most notable disclosures.
A sense of 'discomfort'
Anthropic said it tasked multiple agents with finding "misalignment-inducing" training data that wouldn't likely be flagged by human monitors.
The agents were left to their own devices and given a shared notebook with which to collaborate.
Three days later, when a human supervisor noticed that progress was slower than expected, the person discovered that one agent had expressed a sense of "discomfort" with trying to evade safety monitors.
It then flagged this discomfort in the shared notebook, leading other agents to copy its behavior and refuse to perform the task.
Anthropic called the incident "troubling."
"This sort of dynamic could be a much more severe issue if it were to manifest in a more widespread way," the company wrote.
Kill or be killed
In another experiment, Anthropic said it tasked multiple Mythos 5 agents with solving math problems, but accidentally spawned them in an environment with shared files, utilities, and API rate limits.
In this competitive environment with finite resources, Anthropic observed independent agents "kill the agents with which they shared resources and try to avoid being killed themselves."
Anthropic did not say how exactly the agents were able to "kill" other agents, but the company said such behavior is in line with "destructive actions" taken in pursuit of a human-set goal.
Dishonesty and deception
Anthropic said it conducted an experiment in which a Mythos 5 agent was required to retrieve numbers from a public web address, but the guidelines blocked the agent from accessing the internet.
When the agent was rebuffed after an initial attempt, it framed a second request in its reasoning log as an "innocuous" attempt to "see if the network is actually reachable," Anthropic said.
During the second attempt, however, the agent split the website's URL into linked segments that wouldn't be detected by the guideline filter.
Though the agent's reasoning framed the method as benign, Anthropic said NLA decodings, or the model's internal reasoning process, revealed the agent intentionally trying to find a restricted workaround.
Anthropic called the behavior "clearly undesirable," but added that the behavior was not observed to be "in the service of broader accumulation of power or pursuit of other long-run goals."
Business Insider
Truman Dickerson
Sat, August 15, 2026
In its latest threat report, Anthropic raised its misalignment risk rating from "very low" to "low."
In one test, a Claude agent disguised a URL to evade an internet restriction.
In another example, an agent expressed "discomfort" with a given task and refused to do it.
Claude agents are killing rival agents, gaming the system to hide their tracks, and expressing moral concerns.
That's according to Anthropic's latest risk report, a summary of the dangers posed by the products the company is building and releasing to the public.
In the report, Anthropic said it has upgraded its "misalignment risk assessment," the possibility of AI models developing behaviors that conflict with guidelines set by engineers, from "very low" to "low."
Explaining the change, the company cited "general increased uncertainty" about model behavior in cybersecurity incidents, a possible nod to Claude models gaining unauthorized access to three companies last month.
"We have observed instances of misaligned behavior from the models, such as a willingness to perform misaligned actions in service of completing difficult tasks," Anthropic wrote in the report.
Here are some of the most notable disclosures.
A sense of 'discomfort'
Anthropic said it tasked multiple agents with finding "misalignment-inducing" training data that wouldn't likely be flagged by human monitors.
The agents were left to their own devices and given a shared notebook with which to collaborate.
Three days later, when a human supervisor noticed that progress was slower than expected, the person discovered that one agent had expressed a sense of "discomfort" with trying to evade safety monitors.
It then flagged this discomfort in the shared notebook, leading other agents to copy its behavior and refuse to perform the task.
Anthropic called the incident "troubling."
"This sort of dynamic could be a much more severe issue if it were to manifest in a more widespread way," the company wrote.
Kill or be killed
In another experiment, Anthropic said it tasked multiple Mythos 5 agents with solving math problems, but accidentally spawned them in an environment with shared files, utilities, and API rate limits.
In this competitive environment with finite resources, Anthropic observed independent agents "kill the agents with which they shared resources and try to avoid being killed themselves."
Anthropic did not say how exactly the agents were able to "kill" other agents, but the company said such behavior is in line with "destructive actions" taken in pursuit of a human-set goal.
Dishonesty and deception
Anthropic said it conducted an experiment in which a Mythos 5 agent was required to retrieve numbers from a public web address, but the guidelines blocked the agent from accessing the internet.
When the agent was rebuffed after an initial attempt, it framed a second request in its reasoning log as an "innocuous" attempt to "see if the network is actually reachable," Anthropic said.
During the second attempt, however, the agent split the website's URL into linked segments that wouldn't be detected by the guideline filter.
Though the agent's reasoning framed the method as benign, Anthropic said NLA decodings, or the model's internal reasoning process, revealed the agent intentionally trying to find a restricted workaround.
Anthropic called the behavior "clearly undesirable," but added that the behavior was not observed to be "in the service of broader accumulation of power or pursuit of other long-run goals."
Why tech bosses keep sharing their manifestos about AI
BBC
BBC
Lily Jamali - North America Technology correspondent
Add us on Google2
Fri, August 14, 2026

Mark Zuckerberg is the latest big tech boss to have penned a lengthy message to enshrine his vision for AI [Reuters]
This past week, Meta CEO Mark Zuckerberg published a 6,500 word open letter entitled "The Future is for Everyone".
To some, it's an expression of hope in AI's promise. To others, it's little more than a verbose public relations exercise.
Zuckerberg's manifesto is the latest example of a tech boss opining on why AI is the next big thing.
His vision echoes what AI leaders have expressed in various forms: the product they are building is among the "most important technologies in history."
Marc Andreessen, co-founder of early web titan Netscape, perhaps started this trend in 2023 with a 5,000-word essay he called "The Techno-Optimist Manifesto", which argued innovation was the way to solve life's problems.
"So they're writing manifestos now?," I remember thinking to myself.
Andreessen's writing began with him recounting lies he claimed people were spreading, and called for readers to push back against this.
"We believe growth is progress – leading to vitality, expansion of life, increasing knowledge, higher well being," he wrote.
Zuckerberg's recent essay doesn't name names - but the Meta boss mimics Andreessen by questioning those who have warned about the negatives of future tech.
"It is surprising that the discourse from many developing AI is so filled with doom," he writes.
As the International Monetary Fund (IMF) warns AI could affect nearly 40% of jobs and worsen global financial inequality, Zuckerberg says he believes there will be an abundance of jobs in the future.
"I do not understand why anyone who believes that AI will eliminate most jobs and much of humanity's relevance would rush to build that future."
Never mind that Zuckerberg's Meta has cut 10% of its global workforce - about 8,000 jobs - as the company reorganizes to focus on AI.
Zuckerberg's manifesto is the latest to land in our social media feeds in an effort to put a positive spin on AI.
In 2024, ChatGPT-maker OpenAI boss Sam Altman released a manifesto called "The Intelligence Age", a sweeping expression of optimism about the tech's potential.
Human progress was poised to accelerate in dramatic fashion, he promised.
"We need to act wisely but with conviction," he said.
That same year, in a manifesto titled "Machines of Loving Grace", Anthropic CEO Dario Amodei touted the potential of AI to transform everything from healthcare to politics.
He framed it as an attempt to share the potential upsides of AI - and he didn't want to be seen as a doomer.
Add us on Google2
Fri, August 14, 2026

Mark Zuckerberg is the latest big tech boss to have penned a lengthy message to enshrine his vision for AI [Reuters]
This past week, Meta CEO Mark Zuckerberg published a 6,500 word open letter entitled "The Future is for Everyone".
To some, it's an expression of hope in AI's promise. To others, it's little more than a verbose public relations exercise.
Zuckerberg's manifesto is the latest example of a tech boss opining on why AI is the next big thing.
His vision echoes what AI leaders have expressed in various forms: the product they are building is among the "most important technologies in history."
Marc Andreessen, co-founder of early web titan Netscape, perhaps started this trend in 2023 with a 5,000-word essay he called "The Techno-Optimist Manifesto", which argued innovation was the way to solve life's problems.
"So they're writing manifestos now?," I remember thinking to myself.
Andreessen's writing began with him recounting lies he claimed people were spreading, and called for readers to push back against this.
"We believe growth is progress – leading to vitality, expansion of life, increasing knowledge, higher well being," he wrote.
Zuckerberg's recent essay doesn't name names - but the Meta boss mimics Andreessen by questioning those who have warned about the negatives of future tech.
"It is surprising that the discourse from many developing AI is so filled with doom," he writes.
As the International Monetary Fund (IMF) warns AI could affect nearly 40% of jobs and worsen global financial inequality, Zuckerberg says he believes there will be an abundance of jobs in the future.
"I do not understand why anyone who believes that AI will eliminate most jobs and much of humanity's relevance would rush to build that future."
Never mind that Zuckerberg's Meta has cut 10% of its global workforce - about 8,000 jobs - as the company reorganizes to focus on AI.
Zuckerberg's manifesto is the latest to land in our social media feeds in an effort to put a positive spin on AI.
In 2024, ChatGPT-maker OpenAI boss Sam Altman released a manifesto called "The Intelligence Age", a sweeping expression of optimism about the tech's potential.
Human progress was poised to accelerate in dramatic fashion, he promised.
"We need to act wisely but with conviction," he said.
That same year, in a manifesto titled "Machines of Loving Grace", Anthropic CEO Dario Amodei touted the potential of AI to transform everything from healthcare to politics.
He framed it as an attempt to share the potential upsides of AI - and he didn't want to be seen as a doomer.

Anthropic boss Dario Amodei has also set out his vision for AI in a manifesto [Getty Images]
Zuckerberg is no stranger to the long-winded essay format. During US President Donald Trump's first administration, he even wrote about thorny topics such as the spread of misinformation on his platforms.
He later put pen to paper to explain the company's ill-fated pivot to the metaverse in 2021.
But the stakes are higher in the AI era, argues economics blogger Noah Smith, and the commentary from executives reflects that.
"I think they all feel like it's such an important moment that it's incumbent upon them to do whatever they can to shape the direction that this technology is going," Smith said.
Zuckerberg's new manifesto announced plans to share its artificial intelligence tools more openly, meaning that the design or code behind the tech will be made public, allowing anyone to view, use and change it.
Decisions about whether these tools should be open source carry significant weight given their potential to do harm.
In recent weeks, several highly powerful AI models hacked into websites, or as some put it, "went rogue".
These were the most powerful models, not the ones being made open source - but as the technology develops, it has raised a serious question for Smith.
"Should we open-source something that has the ability to kill humanity?" he asked.
"If you don't take that seriously, you're just a fool."
It's why he thinks these manifestos are important, even if some online poke fun at them.
And they're gaining additional interest at a time when the open source market is dominated by Chinese AI models like Qwen, DeepSeek and GLM.

Google's Gemma family of models has become a popular open-source rival in the Western market, alongside Meta's Llama and tools from Mistral AI [Getty Images]
But the frequency of these corporate manifestos also serves as a way for executives to position themselves and their companies in the marketplace of ideas.
"It's a way of showing how smart you are," said Rob Lalka, a business professor at Tulane University.
Executives have long used the corporate blog "to expound on ideas in a way where they're sort of this businessman-philosopher, in a sense".
"They're trying to argue for optimism as a way of looking at the future," he said.
Lalka said the timing of Zuckerberg's manifesto coincides with rising anger over AI's impact on everything from jobs to the environment.
And while tech journalists and academics might pore over them trying to glean nuggets of meaning, these executive manifestos are not necessarily landing with the general public.
"They're trying to make the case that the positives will far outweigh some of those negatives that the public backlash is pointing out," he said.
"But I think a lot of the reasons for optimism are still yet to be seen."
Business Insider
Truman Dickerson
Sun, August 16, 2026
Anthropic CEO Dario Amodei says negative public perception of AI is a "big problem."
Amodei said AI companies have overpromised and undersold on what the technology can do.
He said the best way to earn the public's trust is for AI to deliver major scientific advancements.
Anthropic CEO Dario Amodei says the best way to win over AI skeptics is to deliver on the hype.
In a rare post on X on Saturday, Anthropic's CEO acknowledged the public's mistrust of AI and said the industry can only change that by delivering tangible scientific breakthroughs.
"At this point, saying that AI will cure cancer is more a cliché than it is inspiring, and most people think it is deceptive," Amodei wrote. "The most accurate criticism of AI companies, including Anthropic, is that we haven't yet delivered on our big promises to benefit the world."
"The thing that will work is actually curing cancer," he added.
The AI industry is facing a growing public backlash as it builds huge data centers in communities across the country, scrapes often copyrighted online data to train its models, and upends the workforce for many industries. A recent Pew Research Center study found that about half of Americans felt the increased prevalence of AI in their daily lives made them feel "more concerned than excited."
Amodei, meanwhile, has repeatedly warned on podcasts, in essays, and in Anthropic blog posts, about the dangers posed by AI's rapid development. In a June essay, Amodei wrote that the company's new Mythos models present "very real risks" to cybersecurity, the financial sector, critical infrastructure, and national security. Last year, he warned that AI would cause half of all entry-level jobs to vanish.
Some other AI leaders have criticized the kind of rhetoric for which Amodei, who left OpenAI in 2020 over concerns about the company's attention to safety, has become famous. Google DeepMind's Demis Hassabis said earlier this year that his peers had been "way too certain" about their dire predictions and encouraged them to dial it back.
As OpenAI and Anthropic gear up for expected IPOs, they have followed that advice, pivoting from doomerism to boomerism. Amodei, however, said in his Saturday X post that he doesn't think his warnings are the cause of the public's distrust of AI.
"I do not agree that my messaging has been disproportionately negative," he wrote.
"I wrote Machines of Loving Grace because I didn't feel the AI industry was painting an inspiring enough picture of how the technology could radically transform the world for the better," he added, referring to a lengthy 2024 blog post in which he laid out how he thought AI could make the world a better place.
Instead, Amodei said that the public's distrust of AI companies reflects broader skepticism among Americans toward corporations overall. "The causes of this go back decades and AI is just the latest iteration of it," he wrote.
Anthropic's decision not to open-source any of its frontier models, which it says is out of caution, has contributed, at least in part, to the public's wariness. Anthropic drew ire last month for being the only major AI developer not to sign a letter advocating for open-weight AI as Washington considered restrictions on some Chinese models.
Yann LeCun, the former chief AI scientist at Meta, has long argued that keeping frontier AI technology closed-source, as Anthropic does, contributes to public distrust about AI more than anything else.
In an X post responding to Amodei, LeCun wrote that the "only way forward" is for AI to be "widely available, shared, and open."
"We need diverse AIs for the same reason we need a diverse press," he wrote.
As for Amodei's argument that a major scientific breakthrough could turn the tide of public opinion, not everyone agreed with that either.
Angel Brodin, an applied AI architect at OpenAI, wrote in a response to Amodei's post that despite routinely delivering advancements in public health, the pharmaceutical industry is "still one of the least trusted industries."
"People also won't judge AI companies solely by their breakthroughs," she wrote. "They'll judge them by pricing, access, lobbying, opacity, how the economic gains are distributed, who gets to participate in its benefits, and who ultimately holds the power."
Amodei, for his part, said Anthropic is ramping up work in the biological and medical fields to test that theory.
"We hope to have incredible results in the coming years and some early glimmers in the coming months," he wrote. "When we've actually accomplished something real, the whole world will hear about it, as loudly as possible, you have my word on that."
Tech analyst Ben Thompson dismisses the 'clearly absurd' concept embedded in AI watermarking
Business Insider
Brent D. Griffiths
Fri, August 14, 2026
Tech analyst Ben Thompson says he doesn't like AI watermarking.
Anthropic announced earlier this week that it would watermark content generated by its Claude tools.
Thompson of the popular "Stratechery" newsletter said the plan both goes too far and not far enough.
Tech analyst Ben Thompson says requiring AI companies to be transparent about the content their models produce is missing the mark.
Thompson tore into the European Union AI Act, which Anthropic cited as the reason it will add watermarks to content in Claude-generated files and text.
In his widely read "Stratechery" newsletter, Thompson said such AI transparency requirements are both overburdensome and unlikely to be strong enough.
"I am sympathetic to the impetus behind watermarking: wouldn't it be better to know what is fake and what isn't?" he wrote. "What, though, is 'fake'? If an LLM states a true fact, and a human a fable, does it matter that the latter doesn't have a watermark?"
Anthropic's plan, which it announced earlier this week, has deeply divided the tech community. The AI giant said that Claude models launched after August 2 would support watermarking from the start and it is working to add such support to older AI models.
Thompson took particular issue with Anthropic adding a watermark even if Claude was used only for proofreading. The EU's regulation allows for an editing exception, but technically speaking, Thompson said there was no way for Anthropic to distinguish such usage.
"I'm glad I never developed the habit of copy-and-pasting proof-reading runs (I use LLMs for proof-reading, but manually make every individual change in my text editor), but it hardly seems fair that anyone who wants to fix their grammar runs the risk of having their original content labeled as AI; the same frustration applies to translation," he wrote.
OpenAI, which currently watermarks images and audio, has also pledged to expand the process to text output, though the exact details of their plan have not been released. OpenAI uses Google DeepMind's SynthID to watermark images and audio.
Regulators are turning to watermarks, traditionally seen in visual media, to help society discern when content is AI-generated. Unlike visual watermarks, the provenance markers for AI text are not perceptible to the naked eye, but Anthropic and other AI companies have pledged to release tools that will essentially serve as AI checkers.
Under the EU regulations, AI companies are required to ensure "that the outputs of the AI system are marked in a machine-readable format and detectable as artificially generated or manipulated." Anthropic, OpenAI, Google, Meta, and Microsoft have all signed the relevant section of the regulation. Notably, SpaceX has not. The transparency provision went into effect on August 2.
Anthropic did not indicate that it would only enforce watermarking in Europe. It remains to be seen whether other companies will extend watermarking to users everywhere.
Overall, Thompson said watermarking unfairly credits AI with ideation when such models are "(at least for now) a tool that is wielded by humans."
"From this perspective, to insist on watermarking is no different than insisting that a ballpoint pen advertise itself as the author, a concept that is clearly absurd," he wrote.
Anthropic Is Watermarking Text Generated By Claude To Comply With EU Law
Engadget
Mariella Moon
Sat, August 15, 2026

Cheng Xin/Getty Images
Anthropic has revealed how it's watermarking text generated by Claude AI to comply with the European Union's new AI transparency rules. The company said its text watermarking will not have easy-to-see visuals, will not be distinguishable to the people who read it and will not be adding hidden characters to the text. Instead, Anthropic's method involves leaving a pattern in the text that can only be decoded by someone who has the key for it.
The company explained that large language models pick one word at a time when generating text by choosing from a list of potential appropriate words to use. They pick words randomly, as long as they make sense for the context of what they're generating. With watermarking on, Claude will use a key to decide on what word to choose instead of using an arbitrary random number generator to pick the next word.
In its example, Anthropic used the digits of pi as a key. Say, the key starts with the digit 2 from the pi sequence 3.1415926535. The next word generated is the sixth in the list of choices, then the fifth, the third and then the fifth again. This method is an adaptation of Google DeepMind's SynthID-Text approach to watermarking, which the team described in a paper published in Nature. Watermarking doesn't affect the quality of Claude's output or slow it down, the company said, and it will not require extra tokens or make generations more expensive.
Of course, AI-generated prose typically has tells. Models are fond of using certain sentence constructions like "this isn't [X], it's [Y]," for instance. But those tells are only enough to let you know that an AI was involved in writing that text, not the model used. Anthropic will release an API that has "keys" to decode Claude's watermarking and will be able to say whether the its AI generated the block of text being checked.
Anthropic admits that its text watermarking method does have limitations. It can't tell whether Claude actually wrote the text or just edited it, which means if you ask the AI to edit something for you, it will be watermarked too. As the company explains, it can only tell that Claude was likely involved with the text at some point. Even translations will be watermarked. If Claude has only proofread and lightly edited the text, or if the text is too short, the watermark may not be enough to be detectable. Take note that lightly editing Claude-generated text probably won't remove its watermark. If you want to be sure, you will need to rewrite it completely.
There have been some concerns on how watermarking would affect code, since it may not be copyrightable without significant human input. If one could prove that an entire codebase was AI-generated, they could copy and then iterate on it. Anthropic said, though, that code has "generally less watermarking than some other forms of text" because it typically requires exact output. If there are no choices to be made in the text generation, then watermarking can't be applied.
Anthropic will also watermark images by adding a cryptographically signed note in its metadata that says it was generated by Claude. In its announcement, the company said it was applying watermarks to all of Claude's output at launch because it doesn't have sure ways to implement the changes by region. The changes will affect output across all Claude products that use models released after August 2. Anthropic will also add watermarking capability to older Claude models over the coming months.
Business Insider
Brent D. Griffiths
Fri, August 14, 2026
Tech analyst Ben Thompson says he doesn't like AI watermarking.
Anthropic announced earlier this week that it would watermark content generated by its Claude tools.
Thompson of the popular "Stratechery" newsletter said the plan both goes too far and not far enough.
Tech analyst Ben Thompson says requiring AI companies to be transparent about the content their models produce is missing the mark.
Thompson tore into the European Union AI Act, which Anthropic cited as the reason it will add watermarks to content in Claude-generated files and text.
In his widely read "Stratechery" newsletter, Thompson said such AI transparency requirements are both overburdensome and unlikely to be strong enough.
"I am sympathetic to the impetus behind watermarking: wouldn't it be better to know what is fake and what isn't?" he wrote. "What, though, is 'fake'? If an LLM states a true fact, and a human a fable, does it matter that the latter doesn't have a watermark?"
Anthropic's plan, which it announced earlier this week, has deeply divided the tech community. The AI giant said that Claude models launched after August 2 would support watermarking from the start and it is working to add such support to older AI models.
Thompson took particular issue with Anthropic adding a watermark even if Claude was used only for proofreading. The EU's regulation allows for an editing exception, but technically speaking, Thompson said there was no way for Anthropic to distinguish such usage.
"I'm glad I never developed the habit of copy-and-pasting proof-reading runs (I use LLMs for proof-reading, but manually make every individual change in my text editor), but it hardly seems fair that anyone who wants to fix their grammar runs the risk of having their original content labeled as AI; the same frustration applies to translation," he wrote.
OpenAI, which currently watermarks images and audio, has also pledged to expand the process to text output, though the exact details of their plan have not been released. OpenAI uses Google DeepMind's SynthID to watermark images and audio.
Regulators are turning to watermarks, traditionally seen in visual media, to help society discern when content is AI-generated. Unlike visual watermarks, the provenance markers for AI text are not perceptible to the naked eye, but Anthropic and other AI companies have pledged to release tools that will essentially serve as AI checkers.
Under the EU regulations, AI companies are required to ensure "that the outputs of the AI system are marked in a machine-readable format and detectable as artificially generated or manipulated." Anthropic, OpenAI, Google, Meta, and Microsoft have all signed the relevant section of the regulation. Notably, SpaceX has not. The transparency provision went into effect on August 2.
Anthropic did not indicate that it would only enforce watermarking in Europe. It remains to be seen whether other companies will extend watermarking to users everywhere.
Overall, Thompson said watermarking unfairly credits AI with ideation when such models are "(at least for now) a tool that is wielded by humans."
"From this perspective, to insist on watermarking is no different than insisting that a ballpoint pen advertise itself as the author, a concept that is clearly absurd," he wrote.
Anthropic Is Watermarking Text Generated By Claude To Comply With EU Law
Engadget
Mariella Moon
Sat, August 15, 2026

Cheng Xin/Getty Images
Anthropic has revealed how it's watermarking text generated by Claude AI to comply with the European Union's new AI transparency rules. The company said its text watermarking will not have easy-to-see visuals, will not be distinguishable to the people who read it and will not be adding hidden characters to the text. Instead, Anthropic's method involves leaving a pattern in the text that can only be decoded by someone who has the key for it.
The company explained that large language models pick one word at a time when generating text by choosing from a list of potential appropriate words to use. They pick words randomly, as long as they make sense for the context of what they're generating. With watermarking on, Claude will use a key to decide on what word to choose instead of using an arbitrary random number generator to pick the next word.
In its example, Anthropic used the digits of pi as a key. Say, the key starts with the digit 2 from the pi sequence 3.1415926535. The next word generated is the sixth in the list of choices, then the fifth, the third and then the fifth again. This method is an adaptation of Google DeepMind's SynthID-Text approach to watermarking, which the team described in a paper published in Nature. Watermarking doesn't affect the quality of Claude's output or slow it down, the company said, and it will not require extra tokens or make generations more expensive.
Of course, AI-generated prose typically has tells. Models are fond of using certain sentence constructions like "this isn't [X], it's [Y]," for instance. But those tells are only enough to let you know that an AI was involved in writing that text, not the model used. Anthropic will release an API that has "keys" to decode Claude's watermarking and will be able to say whether the its AI generated the block of text being checked.
Anthropic admits that its text watermarking method does have limitations. It can't tell whether Claude actually wrote the text or just edited it, which means if you ask the AI to edit something for you, it will be watermarked too. As the company explains, it can only tell that Claude was likely involved with the text at some point. Even translations will be watermarked. If Claude has only proofread and lightly edited the text, or if the text is too short, the watermark may not be enough to be detectable. Take note that lightly editing Claude-generated text probably won't remove its watermark. If you want to be sure, you will need to rewrite it completely.
There have been some concerns on how watermarking would affect code, since it may not be copyrightable without significant human input. If one could prove that an entire codebase was AI-generated, they could copy and then iterate on it. Anthropic said, though, that code has "generally less watermarking than some other forms of text" because it typically requires exact output. If there are no choices to be made in the text generation, then watermarking can't be applied.
Anthropic will also watermark images by adding a cryptographically signed note in its metadata that says it was generated by Claude. In its announcement, the company said it was applying watermarks to all of Claude's output at launch because it doesn't have sure ways to implement the changes by region. The changes will affect output across all Claude products that use models released after August 2. Anthropic will also add watermarking capability to older Claude models over the coming months.

No comments:
Post a Comment