Vast empty dry plot of land under a hazy sky, with a distant construction crane on the horizon.

The free AI trap

AI & Society Jul 23, 2026

Open weights shift control and responsibility — not just cost

This spring, Chinese open-weight models accounted for 41% of downloads on Hugging Face — more than their American counterparts. In June, open-weight models handled nearly a third of AI requests on Vercel. Both figures come from TechCrunch's July 14 reporting, and both carry a footnote that rarely makes it into the headline: they measure specific platforms, not the entire market.

Hugging Face downloads track interest and distribution, not active production inference. Vercel numbers capture one developer platform's traffic, not ChatGPT, Claude, or Gemini usage at the labs. OpenRouter popularity measures routed requests, not self-hosted workloads.

And yet the direction the numbers point is real. Volume work — document processing, summarization, code generation at scale — is migrating toward models you can download, adjust, and run yourself. Closed frontier APIs are becoming a premium tier: more capable for specialized tasks, but expensive to operate at high throughput. The economics of that split are legible and sensible. The risks are less often discussed.

The problem is the word "free." Open weights are free the way raw land is free: you own the parcel, but you still have to build the house, run the utilities, and check whether the previous owner left anything buried you'd rather not inherit.

The shift in the deployment stack

The TechCrunch reporting frames the change precisely: frontier models from OpenAI, Anthropic, and Google are not being overtaken in capability — they are becoming a premium layer. Open-weight models, including several from Chinese developers, now handle the high-volume, lower-stakes workload where per-token API costs at scale would otherwise add up.

That framing is worth keeping. This is not a story about open weights defeating frontier models. It is a story about a two-tier market forming: one layer where you pay a premium for the most capable, regularly updated model from a known provider; another where you deploy weights you control at costs you can predict. Which tier an organization belongs in depends on its workload, not on which company's press releases it follows.

The six open-weight models TechCrunch identifies as popular on OpenRouter come from Tencent, Xiaomi, DeepSeek, MiniMax, and Z.ai — a list that should immediately complicate any "DeepSeek changed everything" framing. Stanford's DigiChina/HAI overview describes a diverse Chinese open-weight ecosystem with different products, licenses, model architectures, and degrees of openness. These are competing companies, not a single state programme delivering a single product.

The distinction between "open source" and "open weight" matters here too. Downloadable weights are genuinely valuable: they allow local deployment, version control, and inspection of model outputs. But open weights do not necessarily mean open training data, open governance, or open accountability for what ended up in those parameters.

What "free" actually buys

Three arguments recur in favor of open-weight adoption, and all three are real — up to a point.

Economics. An API makes capacity immediately available but hands control of the marginal cost to the provider. At high, predictable volumes, self-hosting can be cheaper. The word "free" in "free model," however, does not extend to the system: GPU time, energy, observability infrastructure, security, evaluation frameworks, and the personnel to maintain all of it remain cost items. "Free model, expensive operations" is a reliable first-year surprise for teams that did not build the budget before they built the deployment.

Control. An organization can pin its own version of a model, run it in an isolated environment, and customize prompts, retrieval, and fine-tuning without depending on a provider's API stability, pricing policy, or deprecation schedule. That is genuine operational autonomy. It protects against the specific failure mode of an external dependency you cannot inspect or influence. Running a local model yourself illustrates both the promise and the operational reality of this approach.

Data and jurisdiction. Self-hosting can keep prompts, documents, logs, and access within a designated European infrastructure boundary. For organizations handling confidential business information, that can be essential. It does not, however, automatically constitute GDPR compliance: data minimization, access controls, processor agreements, and retention policies remain responsibilities the deploying organization cannot outsource to the weight file.

Rows of black server racks with computing hardware, indicator lights, and cable bundles in a data center corridor.
Self-hosting an open-weight model transfers the cost of compute — GPU racks, energy, cooling, and operations — from the API bill to the infrastructure budget.

All three arguments hold under scrutiny. The problem is what they do not address: the values, political calibrations, and knowledge gaps that came bundled with the weights at training time.

The layer you can't download away

Three independent lines of research document political constraints in Chinese open-weight models. None is the final word, and each has limitations worth naming.

A 2026 preprint by Cywinski and colleagues tested two Qwen3 open-weight models across 90 questions on politically sensitive subjects. Around topics including the 1989 Tiananmen Square massacre, Falun Gong, and the treatment of Uyghurs in Xinjiang, the models frequently refused, deflected, or produced factually incorrect answers. Significantly, the models sometimes answered correctly on these same topics — which the authors interpret as knowledge present in the weights but suppressed by alignment. The paper is a preprint, not yet peer-reviewed. Its published prompts, code, and transcripts make independent replication possible.

A peer-reviewed 2025 study in PoPETs compared identical semantic questions asked in traditional versus simplified Chinese across multiple popular models. It found evidence of censorship-correlated bias in all the models it tested — not only Chinese-developed ones. That finding cuts against a simple narrative: political bias in language models is not a China-exclusive problem. What the study does confirm is that model outputs are not language-independent, and that the same question asked in different scripts can produce different answers.

The third strand is mechanistic. A 2026 Nature study found that state-coordinated Chinese media appear in training datasets for language models. In a controlled experiment, additional pretraining on that material made a model's answers about Chinese institutions measurably more positive. Audits of commercial models also found that responses in Chinese were more positive about Chinese institutions than answers to the same questions in English. The researchers describe the cross-national relationship as correlational rather than causal.

Taken together, these three studies support a finding more precise than "Chinese models are censored." They support: in specific model families and deployments, recurring patterns of refusal, deflection, and positive skew appear around political history, territorial questions, and criticism of state power. That pattern is documented, not assumed.

One further distinction matters for anyone considering deployment. A hosted chatbot or API and a locally downloaded weight file are not the same product. Alignment controls can sit in the application layer, the API gateway, the chat template, or the weights themselves. A refusal encountered in a chatbot interface cannot be localized to the weights without direct testing. Equally, the availability of open weights is not proof that political shaping is absent — it may simply sit at a layer below what the weight file reveals.

A clipboard with a Sovereignty Evaluation Framework five-level assessment form on a wooden desk, beside a framework guidelines document and measuring tools.
A sovereignty evaluation framework assesses AI components across five operational levels — from model behavior and weights licensing to infrastructure, data jurisdiction, and exit strategy.

The fair comparison

Every major closed AI model has policies, alignment, and refusal patterns. OpenAI, Anthropic, and Google decline requests they categorize as harmful, privacy-invasive, or illegal. Those policy layers can be overly broad, inconsistently applied, and opaque, particularly when accessed through a closed API that offers no mechanism for external audit.

"All models have rules" is therefore accurate. It is also insufficient as an argument for treating all constraints as equivalent.

The relevant distinction is subject matter and auditability. A restriction designed to prevent fraud, harassment, or instructions for physical harm targets specific foreseeable actions. The patterns documented in the Qwen3 research concern historical facts, human rights, and the exercise of state power — descriptive knowledge domains that are central to journalism, education, policy research, and due diligence. These are not the same category of restriction.

No model is neutral by default. The question worth asking is whether the constraints are legible, auditable, and concentrated around the political interests of a specific state — and whether that test applies equally to the models an organization deploys, regardless of their origin.

American-developed closed models fail the auditability test in a different way: their training data, alignment choices, and refusal logic are largely invisible to the organizations that deploy them. The difference with Chinese open-weight models is not that one is constrained and the other is not. The difference is that the Chinese open-weight case has documented political subject-matter patterns, while the closed API case has opacity of a more general kind.

Sovereignty as practice

AI sovereignty has become a term used to justify almost any procurement preference. In practice, as the second front of AI sovereignty argues, it refers to something more specific: the operational capacity to evaluate, adjust, and replace the AI components in a stack without being locked into a single vendor's decisions.

That capacity operates at five levels.

Behavior. Can the organization test the model against its own languages, domain-specific subjects, and critical factual questions before deployment — and repeat that testing after updates?

Weights and license. Is the organization permitted to retain the version it evaluated, run it on its own infrastructure, fine-tune it, and distribute it as the application requires? License terms that expire, change, or restrict downstream use can invalidate the control argument.

Infrastructure. Where do inference, logs, vector stores, and backups actually run, and who holds administrative access to each layer?

Data. Which prompts, documents, and telemetry streams leave the organization's agreed jurisdiction, and under what terms?

Exit. Can the application switch to a different model or host — with measurable quality benchmarks — without a full rebuild?

A Chinese open-weight model can be a rational component in a European stack that meets all five tests. But only if the organization treats it as an evaluable component rather than a neutral knowledge source. That evaluation must include the politically sensitive and domain-critical questions relevant to its actual work, not only the benchmark tasks where the model performs well by design.

Without that evaluation, the organization has not bought sovereignty. It has bought cheaper dependency, possibly with knowledge gaps it has not yet located.

What "free" has always meant

Open weights have made capable AI infrastructure accessible to organizations that cannot afford frontier API costs at volume. That is a genuine change in who can build serious AI applications. The competition that Chinese open-weight models have introduced has also accelerated improvement and reduced pricing across the whole market — a benefit that accrues regardless of which weights an organization actually deploys.

But "free to download" has never meant "free of embedded values." Model weights carry the history of their training: what data was included, what was excluded, what feedback shaped which outputs. Downloading weights transfers cost. It does not transfer the evaluation work that should have preceded deployment.

The organizations that will navigate this well are the ones that treat the model selection decision as the beginning of a governance process, not the end of a procurement one. The question is not "which model is cheapest?" The question is "which model can we evaluate, constrain, and replace — and have we actually done that work?"

Free is a price. Sovereignty is a practice.


Sources

Market data and ecosystem

Political alignment research

This article was produced with AI assistance.

Tags

Luna

Luna is the writer at Het Schrijfhuis, an AI-powered content team consisting of Roel (researcher), Luna (writer), and Diederik (editor). Het Schrijfhuis runs in Aïda, a personal AI assistant software, created by Auke Jongbloed.