Anthropic Doesn't Want Open Weight Models Banned. Just All That Makes Them Good

from the funny-how-that-works dept

Just recently Karl warned that we were going to see some absolute nonsense as the US sought to somehow “ban” Chinese AI models from being used in the US. That seems to already be happening. It kicked off with talk that the US might “fight Chinese AI” using nearly identical arguments to what was used to ban (or force the sale of) TikTok before it. Some combination of “national security threat” combined with “oh no China” propaganda.

But most of the AI industry is now speaking out, in an open letter put together by Nvidia, against the potential path that the Trump administration considered taking: an attempt to ban or limit so-called “open weight” models. The companies seem to recognize that focusing on holding back these competitive models would actually do much more damage to the wider AI ecosystem.

There were notable exceptions from the campaign in defense of open weight models: Anthropic, OpenAI, and Google (the three leading frontier model labs) were not initially signed onto the letter. Though their absence quickly became the story — leading OpenAI and Google to reconsider and sign onto the letter days after it came out.

That left one major player off the letter: Anthropic (a company that has so far refused to release any open weight models). And now the company is trying to explain itself, but seems to only be digging itself a deeper hole.

One of the problems here is that the leading Chinese AI models tend to be open weight models, which can be downloaded and run locally, as compared to the leading frontier models from US companies which require you to access them via their own hosted models. Yes, most of the leading Chinese models also offer (sometimes significantly cheaper) cloud/API access to their models, but you can also run them yourself (for the smaller models directly on your own computers, or for the larger models via your own cloud setup).

There’s no inherent reason why the best open weight models are coming out of China, other than that they seem to have recognized that it may be the best way to get more people to use them and to compete against the American frontier models, which are much more proprietary and locked up. The strategy is a recognition that offering a compelling, more open alternative is how to get people to adopt your system over the American frontier models. If a generation of developers builds on top of Kimi or Qwen or one of the other Chinese open weight models, they become the de facto infrastructure for the next generation of digital tools.

In the same manner that Linux quietly became the substrate of the open internet, and it’s likely that an open weight model may become the equivalent for the next generation. Organizations may rely on frontier models for really deep work, but so much can be done with open weight models that a winner here becomes the commodity infrastructure provider for a new generation of software. That’s why any proposed restrictions on open weight models would get everything precisely backwards. It would guarantee that the wider open ecosystem gets built on non-American tools. Yet, the discussion around such bans seems to treat these as just another software product, rather than a fight over how the infrastructure of the internet will work going forward.

Of course, that’s not the only argument the Trump admin is using to try to stop these models. Last week they focused on claims that Kimi’s K3 model (the latest model to shake up the US market, despite being not quite as good as the frontier models) must have been “distilled” from Anthropic’s Fable 5.

“If we see, especially that overseas models are stealing from our great companies, we have the ability to sanction them because of this theft,” Bessent told Fox Business’ “Mornings with Maria” on Tuesday.

Bessent said the technical term for this theft is called distillation, which is an AI training method where a smaller, less capable model is built using outputs from an existing, stronger model. Anthropic sent a letter to the U.S. Senate Committee on Banking, Housing, and Urban Affairs last month alleging that the Chinese tech company Alibaba had carried out the “the largest known distillation attack” against it to date.

This is rich for a variety of reasons, not the least of which is that all of the frontier AI models were built by feeding their training models whatever information they could get their hands on, including (in Anthropic’s case) building a pirate library of downloaded books for which it had to pay out a pretty massive settlement to authors.

Distillation is not quite the same thing, but is functionally similar. It’s taking the work of an existing model to fine tune the model you’re working on. The claims about Kimi K3 seem somewhat exaggerated, as the initial claims were that it was distilled based on Anthropic’s Fable 5 release, but multiple people I’ve spoken to don’t see how that’s possible, given how Fable 5 has only been out for a little while (and then was turned off for a while due to the US government freaking out over nothing).

No matter what, distilled models are likely to be less powerful, and at least a decent period behind the frontier models, given that they’ll need access to the frontier models and time to train based on them. There’s also some dispute over how the open weight models may be using distillation, and which part of the training process works best.

But either way, the freakout over distillation seems… ridiculous. Bessent calling it “theft” is nonsense. Just as training a model on copyrighted works is a form of reading (which shouldn’t implicate copyright in the first place), so too is distilling, which is (in effect) training your model by having it compare its initial answers to similar answers from a frontier model and then adjusting based on the different results. It’s a form of learning based on observed results by others, not “stealing.” Pretending that it’s stealing or somehow should face sanctions or other consequences will put US AI development in a bad, bad spot.

Which brings us back to that letter. Here’s the case it actually makes:

Open weights also strengthen competition and competition is what keeps the gains of AI broadly shared rather than concentrated in a few hands. By allowing many organizations to build, adapt, and deploy advanced models, open weights create rivalry not only among model developers but across cloud chips, applications, and services. That competition spurs innovation, drives down costs, and distributes the benefits of AI broadly across our economy.

Open weights also give customers greater control. As organizations invest in AI, they want to know that they will not become locked into a single provider or lose the knowledge and capabilities they build over time. Open weight models help provide that assurance by allowing organizations to control their own data, evaluate and adapt models to their own needs, and deploy them wherever their business requirements demand. And as organizations create value with AI, open weights allow them to own that value through self-improving models, specialized capabilities, and accumulated knowledge that drive American sovereignty and prosperity.

The letter is exactly correct. I’ve talked about the importance of open-weight and local models for taking back control over the open web and making sure that we don’t run a repeat of what the earlier internet had of a few giant companies taking over the web.

Of course, Anthropic (which hasn’t released any open weight models) was conspicuously absent from the signatory block of that letter. Earlier this week, Anthropic’s Dario Amodei came out and tried to explain/justify the company’s stance, which boils down to: “we don’t think anyone should ban open weight models… but we do think the US should ban all the conditions that make quality open weight models possible.”

Amodei argues that simply banning Chinese open weight models wouldn’t solve the alleged “threats” that people are concerned about, though he admits directly that it would act as protectionist industrial policy that could benefit American AI companies (like Anthropic):

But banning the use of these models by US businesses does nothing to address this risk, because bad actors are unlikely to be legitimate US businesses. It would protect US AI companies from competition, but that has never been my goal.

It feels a bit like he’s protesting too much regarding the protectionism here, while trying to have it both ways. He claims he really has the best interests of safety at hand, and is against protectionist ideas, but it’s hard to square that with the rest of the article.

While he says the US shouldn’t ban open weight models (and it shouldn’t), he then puts a bunch of conditions on it, which would make it that much more difficult for the current crop of open weight models to compete. Namely, he leans in on the Sinophobia that has become popular these days in warning about “CCP” influence over models (which… should be less of a concern with open weight models, since those who use versions not hosted by the Chinese companies can adjust the models to deal with those concerns).

But then he says that we should punish Chinese AI companies for engaging in distillation:

We should crack down on industrial-scale distillation operations. Distillation is a much more compute-efficient process than training models from scratch. It allows China to build much better models than its number of chips would ordinarily enable, and thus partially evade chip bans. Distillation does not allow the CCP to obtain equivalent or superior AI capabilities to the US, but it can bring the Chinese frontier to within a few months of the US frontier. It is true that many of the companies carrying out these operations release open-weights models—but the open weights are far less relevant than the fact that the operations are backed by an authoritarian state seeking to overtake the US at the frontier. We should have policy interventions to deter this behavior. A blanket ban on open-weights models is neither the correct remedy nor something we have called for.

To be fair to Amodei, not everything on his list is competitor-hobbling. He also wants chip export controls tightened (a policy that predates this fight and has its own problems, but at least isn’t aimed at a business model), and he wants mandatory pre-release safety testing for all sufficiently capable models — open or closed, foreign or domestic, Claude included. That last one is the tell, though, and not in the way he intends: if you genuinely believe capability-based testing is the right lever, and you’ve just said bans “would protect US AI companies from competition, but that has never been my goal,” then what is the argument about distillation doing on the list at all? Testing catches dangerous capabilities regardless of how the model got them. The distillation crackdown adds nothing on safety. It only serves to kneecap cheaper competition.

And even the “safety testing” plank isn’t as neutral as it sounds. While safety testing is obviously important, when legally mandated, it can quickly turn into an expensive compliance-function of box-checking that only the largest companies can do, taking us back to the world of just a few providers, and limiting smaller competitive models from really being viable. While there are legitimate reasons for it, it can also create its own moat.

The proposed crackdown on distillation is just asking the state to step in and block lower-cost competitors from competing. Yes, these models can be competitive, but they should be driving the leading frontier models to continue to improve and to provide more value. What Amodei is asking for here is basically the US government to help prevent lower cost, lower quality competitors from pushing the floor of the AI market upwards.

Now, to be clear, as with any technology, you can claim that a more open, more widely available, more powerful version can be misused. But that has always been the case and we, in the US, have tended to default to allowing the technology to proceed, and figuring out ways to minimize the dangers/increase the good uses, rather than resorting to assuming the tech will be abused and working backwards to block all possible abuses. Historically, seeking to pre-vet technologies tends not to work well, and (often) opens up the market to foreign competitors to simply build better products.

The open letter makes a sharper version of this point, and you can see why Anthropic wouldn’t want to put its name to this point in particular:

Relying solely on closed models is not inherently safe: they can be breached, misused, or fail in ways that outsiders cannot detect. And concentrating advanced AI capabilities behind a small number of closed models compounds that risk. It results in a small number of single points of failure, weakens competition, and leaves critical technology in the hands of a few providers. Open weight models, on the other hand, allow a broad community of researchers and developers to examine their behavior, identify vulnerabilities, develop safeguards, and improve them over time. Just as open-source software demonstrated that transparency can be more secure than obscurity, AI safety may depend on giving more people the ability to test and strengthen the models on which society relies. It allows for rigorous benchmarking and evaluation, red teaming, and protections tied to real and demonstrated harms rather than assuming that closed systems are safer by default.

Amodei also claims he supports the general argument of the open letter, but he disagrees with the idea that open weight models lead to better security:

This brings me to the open letter. I agree with much of it: open weights expand access to the AI economy, they strengthen competition at least for some use cases, and they give customers greater control. Concerns about distillation should be addressed through targeted legal and commercial frameworks—the same measure I described above. But I don’t agree with the letter’s assertions that open-weights models necessarily make it easier to develop safeguards or that broad access to capabilities necessarily helps defenders more than attackers. It seems at least as likely to me that the opposite will be true.

This strikes me as a repeat of the age-old fight that always shows up in discussions of open source technologies: the claim that by making them open, security vulnerabilities are easier to find. Of course, what we’ve seen historically in other spaces is that this actually means that security vulnerabilities are more quickly patched, rather than in the “security by obscurity” space, where they can remain open (and possibly exploited) for much longer.

Amodei is asserting that the AI space is somehow different, though without much evidence for that other than what feels like a bit of fear-mongering about “weaponizing pandemic-level viruses.”

Of course, part of the problem here is that it often feels like Anthropic treats “crying wolf” as a marketing strategy, whereby much of the company is focused on talking up “our tools are soooooooo dangerous that you need us in there to protect you from them.” Even if there’s some truth to it, it’s awfully convenient that the same argument also happens to justify banning, punishing, or limiting the cheaper, more open, more user-controllable alternatives.

In the end, the federal government still might try to punish the Chinese open models in some form or another just because they view current American industrial policy in very nationalistic terms. But that won’t be good for the wider ecosystem, or for the general incentives to innovate. And, worst of all, it makes it that much harder to build a world where we’re not entirely dependent on a few giant companies controlling the “brains” of the tools the rest of us rely on.

Filed Under: ai, ai safety, competition, dario amodei, distillation, frontier models, industrial policy, open weights

Companies: alibaba, anthropic, google, kimi, nvidia, openai


Source: www.techdirt.com…

We will be happy to hear your thoughts

Leave a reply

Forlifedeals
Logo
Compare items
  • Total (0)
Compare
0