The model is free, the hardware isn't: where value moves when intelligence becomes a commodity
- AI
- Local-AI
- Open-Weight
- Hardware
- Digital-Sovereignty
- Privacy
- GDPR
- Career
On 30 July 2026 OpenAI cut the price of one of its models by 80% in a single day — and not out of generosity: it had launched it just three weeks earlier. The following day, a Chinese lab released a model with comparable performance, weights free to download, at a fraction of the cost.
This is not a price war. It is the signal that the goods have stopped being goods.
For three years we took one thing for granted: that intelligence was the hard part, the scarce part, the part you pay for. That is no longer true. Intelligence has become abundant, replicable and nearly free, while the piece of silicon required to run it has become scarce, contested and increasingly expensive.
This reversal moves value from one end of the supply chain to the other. And it moves, as a consequence, where it makes sense to invest your own time.
First fact: intelligence has become a commodity
Within two months, models that until yesterday would have counted as “frontier” shipped with public, downloadable weights:
| Model | Lab | Parameters | Licence / weights | Note |
|---|---|---|---|---|
| GLM-5.2 | Z.ai (formerly Zhipu), 🇨🇳 | 743B (MoE) | MIT | 51 on the Artificial Analysis Intelligence Index — highest of any open model |
| MiniMax M3 | MiniMax, 🇨🇳 | MoE, 1M context | open | First open model to combine frontier coding with native multimodality |
| Kimi K3 | Moonshot, 🇨🇳 | 2.8 trillion (≈50B active) | open | 1.4–1.5 TB of weights, 47-page technical report included |
| DeepSeek V4 Flash | DeepSeek, 🇨🇳 | 284B | open | Beats its own house’s “Pro” version on agentic benchmarks |
| Inkling | Thinking Machines, 🇺🇸 | 975B (41B active) | Apache 2.0 | First model from Mira Murati, former CTO of OpenAI |
| MAI (7 models) | Microsoft, 🇺🇸 | various | open | The training recipe published too, hyperparameters included |
The point is not any single model. It is the cadence: one a week, for months.
And that cadence produced exactly what you would expect when a good becomes abundant. The price collapsed:
- GPT-5.6 Luna went from $1.00 to $0.20 per million input tokens, and from $6.00 to $1.20 on output. Minus 80%, three weeks after launch.
- DeepSeek V4 Flash sits at $0.14 on input. With cache hits, $0.0028.
- Open Chinese models cost, on average, 60% to 90% less than the flagship offerings from OpenAI and Anthropic.
And developers noticed. A CNBC investigation in July found that Chinese-origin models reached 46% of the tokens consumed by US enterprises on OpenRouter in a single week. Twelve months earlier they averaged 11%; in the first half of 2025, 4.5%.
When two products do the same thing and one costs a tenth of the other, you are no longer selling a product. You are selling a commodity.
Second fact: nobody is making money on tokens
Here is the part that changes the perspective. If intelligence were a healthy business, collapsing prices would be a margin problem. The point is that the margin was not there to begin with.
The trackers comparing AI spend against AI revenue tell one story. The estimates disagree with each other and should be read for what they are — estimates — but the order of magnitude is unambiguous:
- Amazon has reportedly burned hundreds of billions on AI infrastructure against AI revenue in the low tens;
- Meta has reportedly spent over 200 billion to earn back single digits;
- Microsoft and Alphabet sit in the same quadrant.
The hyperscalers collectively are heading toward $700–900 billion in capex in 2026 alone, up more than a third year on year.
The divergence between AI investment and AI revenue currently runs at around 46%. In 2001, at the peak of the telecom excess — the one that preceded a brutal, multi-year correction — it was 32%.
There is, however, one group sitting on the right side of that ledger. Exactly one:
Whoever sells the hardware.
Nvidia and Micron. Not those training the models, not those selling subscriptions, not those building the apps. Those manufacturing the accelerators and the memory.
That single line of the balance sheet retroactively explains almost everything else.
Why everyone suddenly started giving models away
On 24 July 2026, twenty-five American companies signed a letter to Washington — Open Weights and American AI Leadership — asking it not to impose “premature restrictions” on downloadable models, at a moment when the administration was weighing a ban on Chinese models.
The signatories: Nvidia, Microsoft, Meta, IBM, Dell, Palantir, Andreessen Horowitz, Hugging Face, Y Combinator, Mozilla, Mistral, Replit, Perplexity, the Linux Foundation. Initially absent: the two preparing multi-billion-dollar listings, OpenAI and Anthropic. (OpenAI added its name two days later.)
That list is worth reading through business models rather than statements of intent:
- Nvidia releases open models and is their biggest sponsor. Nvidia does not sell tokens: it sells boards. Every additional open model is one more reason to buy hardware.
- Microsoft publishes seven models and even the recipe to train them. Microsoft does not sell tokens: it sells Azure.
- Meta opened Llama years ago. Meta does not sell tokens: it sells advertising, and it needs the ecosystem not to belong to a competitor.
- The Chinese labs give frontier models away. China does not (yet) hold a monopoly on compute — but it holds the factories, and it is building its own chip supply chain. Huawei is targeting roughly 600,000 Ascend 910C units in 2026, and trillion-parameter models pre-trained entirely on domestic clusters already exist.
The pattern is as old as computing: when you cannot monetise a layer, you commoditise it in order to sell the one underneath. Google did it with Android, IBM with Linux, Netscape with the browser.
Except this time the layer underneath is not a cloud service. It is a physical object. And that is where the story gets uncomfortable.
The paradox: software depreciates, hardware doubles
While models become free, RAM has become a luxury good.
Samsung, SK Hynix and Micron control over 95% of global DRAM production, and have deliberately diverted manufacturing capacity toward HBM for AI accelerators. The result: DRAM prices rose roughly 90% in a single quarter, Gartner estimated peaks in the region of 130%, and SK Hynix has warned the shortage could stretch beyond 2030.
The consequences are already on the shelves:
- Dell, HP and Lenovo have raised PC prices by 15–20%.
- On 25 June 2026 Apple raised prices, effective immediately, across all iPads, Macs, HomePods, Vision Pro and Apple TV.
- On 28 July Apple launched Apple Upgrade: a leasing programme, underwritten by Klarna, letting you rent an iPhone, iPad, Watch or Mac over 24–36 months instead of buying it.
Let me pause on that last point, because it is the one I care about most.
We spent fifteen years renting software: streaming, cloud, SaaS, subscriptions. I wrote about it at length when I decided to go back to owning my data with a home server. Now the same logic reaches the physical object. No longer just “you don’t own the films you watch”: soon, potentially, “you don’t own the computer you work on”.
If that is the trajectory — increasingly expensive local-AI machines, sold in instalments — then the question “do I want to own or rent my hardware?” stops being philosophical and becomes a budget line.
Third fact: running a frontier model at home has become genuinely possible
And here, for once, the news is good.
Until a year ago “local AI” meant 7-billion-parameter toys that got arithmetic wrong. Today it means something else, for two precise technical reasons.
1. Mixture-of-Experts architectures broke the equation between size and cost. Kimi K3 has 2.8 trillion total parameters but activates roughly 50 billion per token: only 16 experts out of 896 fire at a time. The model is enormous to store, but relatively cheap to run.
2. SSD streaming broke the RAM wall. Instead of loading the whole model into memory, you stream it in blocks from disk: run a layer, load the next. RAM stops being a binary threshold — “it fits or it doesn’t” — and becomes a speed dial.
The best example of that second idea carries an Italian signature. Salvatore Sanfilippo — antirez, the creator of Redis — has released ds4 (DwarfStar 4): an open source (MIT) inference engine written in pure C that runs DeepSeek V4 Flash locally on 96–128 GB of RAM via asymmetric 2-bit quantisation, and on smaller machines via SSD streaming. It supports Metal, CUDA and ROCm, exposes OpenAI- and Anthropic-compatible APIs, does tool calling and speculative decoding. It passed ten thousand GitHub stars within days.
A near-frontier 284-billion-parameter model. On a laptop. Without sending a single byte to anyone.
At the other end of the spectrum, the market is gearing up: Nvidia’s DGX Station for Windows puts a GB300 Grace Blackwell Ultra on your desk with 748 GB of coherent memory (252 GB HBM3e + 496 GB LPDDR5X) and 20 FP4 petaflops — enough for trillion-parameter models, locally, with no cloud. It reaches market through Dell, HP, ASUS and others in Q4 2026.
And there is talk — a Bloomberg report, unconfirmed by Apple — of a future M7 Ultra with up to 1.5 TB of unified memory, with Blackwell-class AI performance. 2028 horizon, and contingent on the memory crisis easing.
Fourth fact: the model is no longer the variable that matters
There is one data point that I consider the most underrated of all of 2026.
OpenAI published GPT-5.6 Sol’s results on ARC-AGI-3, a benchmark of 2D games where the agent has to infer the rules on its own. With the official harness, the model scored 13.3%. Turning on two API settings — retained reasoning (don’t throw away the model’s reasoning between actions) and compaction (summarise old context instead of truncating it) — the exact same model reached 38.3%. Using six times fewer output tokens.
No retraining. No new model. Just decent working-memory management.
Three times the performance, at a sixth of the cost, by changing the scaffolding around the model.
That is the real message. If raw intelligence is a commodity anyone can download, competitive advantage moves to everything surrounding it: how you manage context, which tools you give it, how you orchestrate the work loops, where you keep the data, what hardware you run it on.
Alibaba pushed this logic to its limit: it claims Qwen3.8-Max (2.4 trillion parameters, 95 billion active) worked 16 consecutive days autonomously, starting from an empty repository and building its own scaffolding along the way — 265 commits, 127 pull requests, 151 issues.
On this one, though, keep your guard up, and I say so gladly because it is exactly the kind of announcement that deserves scepticism. As I write, Qwen3.8-Max’s weights have not been published: the announcement said open weights, but there is nothing on Hugging Face and no licence has been disclosed. One commentator put it well: “an API business model wearing an open source jacket for the launch photo”.
It is a useful reminder: “open” has become a marketing term, and needs verifying case by case. Downloadable weights do not mean a permissive licence, and a permissive licence does not mean traceable training data.
What this means, concretely, for people who work
Put the pieces together. Intelligence costs almost nothing and downloads. Hardware costs more and more and is contested. Value sits in the scaffolding and in where the data lives.
If that picture holds, there is one skill worth building right now: knowing how to run, integrate and secure models on infrastructure the client controls.
This is not an ideological bet. It is a regulatory and contractual one, and the numbers support it:
- On 2 August 2026 the European AI Act became fully applicable. Stacked on top of GDPR, the question “where does this data physically end up?” stopped being a curiosity and started being a condition of signature.
- The sovereign cloud market is worth roughly $80 billion in 2026, growing 35.6% year on year. GAIA-X counts over 400 certified providers across Europe.
- 44% of organisations cite data privacy and security as the primary obstacle to LLM adoption (Kong survey).
- BLS 2024–2034 projections put data scientists at +34% (≈23,400 openings a year) and information security analysts at +29% (≈16,000 a year).
- In regulated environments the preference is explicit: at a Swiss financial industry event, nearly half of the leaders present said they wanted AI processing and storage exclusively within national borders.
A bank, a hospital, a law firm, an R&D department: they cannot ship their most sensitive data to a server whose jurisdiction they do not know. Not out of paranoia — out of contract.
And there is a genuine supply gap here. Almost everyone knows how to use ChatGPT or Claude. Far fewer know how to deploy a model, size it to the available hardware, quantise it without ruining it, build a RAG on top that reads company documents without letting them leave the perimeter, and monitor it.
The good news is that the entry barrier is lower than it looks. If you can use Docker, you are halfway there. ds4, Ollama, llama.cpp and vLLM are one docker run away, and the jump from “I run a small model” to “I run a serious model” today is a question of RAM and patience, not of a PhD.
The third way: don’t move the model, remove the data
There is, however, a practical problem that makes the choice less binary than I have presented it so far.
A six-person law firm, an accountant, a small clinic: they have exactly a bank’s confidentiality problem, but neither the budget for a ten-thousand-euro workstation nor anyone to administer it. And models small enough to run on an ordinary laptop lack the reasoning capacity to analyse a contract or an expert report.
It looks like a dead end. But it is only one if you accept the implicit premise: either you send the document to the model, or you bring the model to the document.
There is a third option, and it is the one I find conceptually the most elegant in this entire picture: leave the reasoning in the cloud and strip the sensitive data before it leaves.
That is the idea behind rizzo-pii, an open source (MIT) project by Simone Rizzo — the same analyst whose work prompted much of this piece. It works in four steps:
- Local detection. A 0.3-billion-parameter token-classification model runs on the user’s own machine and identifies the sensitive entities in the text.
- Deterministic substitution. Each entity is replaced by a typed, stable placeholder —
[FULLNAME_1],[CF_1],[IBAN_1]. Stability is the detail that makes the whole thing work: if the same name appears ten times, it gets the same placeholder ten times, so the remote model still understands who does what to whom. The real mapping goes into an encrypted dictionary that stays on the local disk. - Sending the pseudonymised text. ChatGPT, Claude or Gemini receives a document with its structure intact and its identities removed.
- Local restoration. The response comes back with the placeholders, and the local application swaps the real values back in.
The technical numbers are what make it genuinely practicable: roughly 0.5–1.2 GB of RAM, 64-bit CPU, no GPU, no API key, no network connection. It runs on an office laptop.
And there is a design choice worth noting, because it is the kind of detail that separates a tool built for real use from a demo. The model recognises 22 categories, of which five are specific to the Italian context and are systematically left in the clear by English-centric detectors: tax code, VAT number, cadastral data (sheet, parcel, sub-parcel), act identifiers and province abbreviations. More importantly, for identifiers with a known mathematical structure the neural model does not get the last word: a deterministic verifier checks the checksum — Luhn for cards, mod-97 for IBANs, the official algorithm for tax and VAT codes — and if the code is mathematically valid, the algorithm’s verdict overrides the network’s prediction.
That is the fix for the classic failure of purely neural recognisers: splitting a long code in half and leaving the second portion in the clear. On validation — 7,000 real sentences from Italian legal and administrative texts — the result is a Micro-F1 of 0.989, with 1.000 across all five Italian categories.
A legal distinction worth not skipping
Here, though, rigour matters, because this is the point where it is easy to tell yourself a comfortable story.
What rizzo-pii performs is pseudonymisation, not anonymisation. The distinction is not academic: it is Article 4(5) of the GDPR. Because the local dictionary allows data subjects to be re-identified, the pseudonymised text remains legally personal data under Recital 26.
Three operational consequences:
- You remain the data controller. The tool drastically reduces the risk surface; it does not transfer responsibility to anyone.
- The mapping dictionary becomes a critical asset in its own right. If it accidentally ends up in a folder synced to an unencrypted cloud, you have recreated the problem you were solving — concentrated into a single file.
- Documents processed this way cannot be published as open data, unless you permanently destroy the dictionary and verify that the residual context does not allow indirect re-identification.
That said, the trade-off remains excellent: you keep state-of-the-art reasoning, you eliminate the transfer of identities to third parties, and the entry cost is an office computer rather than a ten-thousand-euro workstation.
It is also, I believe, the architectural pattern most likely to spread: small, specialised local components acting as a compliance filter in front of a large, generalist remote model. Not “all cloud”, not “all in-house”, but a boundary drawn deliberately — and drawn exactly where the data that cannot leave passes through.
Honestly: why local AI is still inconvenient today
I do not want to sell this better than it is. If you go in this direction, here is the bill:
They are slower. A model streaming from SSD on a laptop does not compete with a data centre. Fine for asynchronous work, much less so for an interactive chat.
They are somewhat dumber. Quantising to 2–3 bits costs something in quality. Not a great deal, but something. And open models remain on average a step below the best closed ones: 51 against 56 on the Artificial Analysis index is a real gap, even if far smaller than a year ago.
They cost up front. Local AI shifts spending from variable cost (you pay per token) to fixed cost (you buy the machine). At low volumes the cloud is cheaper, full stop. And in the middle of a memory crisis, this is a terrible moment to buy RAM.
They need maintenance. Updates, backups, security, monitoring. It is not “install and forget”.
Which is why the sensible answer is not “all local”. It is hybrid, with a clear criterion across three tiers:
- Cloud for heavy, exploratory, non-sensitive work — where you need brute force and the data isn’t hot.
- Local for anything that must not leave the perimeter, for predictable and repetitive workloads, and for everything you don’t want a vendor to be able to reprice, re-license or discontinue overnight.
- A local filter in front of the cloud for the most common case of all: when you need the best reasoning available, but on documents containing names, tax codes and bank details.
It was never “cloud bad, local good”. The point is to decide deliberately what to rent and what to own, instead of renting everything by inertia.
Conclusion
For three years the question has been: which model is the smartest?
That question is running out. Not because we solved it, but because the answer has become “more or less all of them, and it changes weekly anyway”. When the hard part becomes free, value migrates: toward hardware, toward scaffolding, toward the jurisdiction the data lives in.
I find it interesting that the extremes of this story touch. On one end, Nvidia builds a 748 GB machine to put a supercomputer on a desk. On the other, two Italian projects, both given away: an engine in C that runs a 284-billion-parameter model on a laptop, and a 0.3-billion model that fits in half a gigabyte of RAM and stops a tax code from landing on an American server.
They are saying the same thing from opposite ends of the scale: control is coming home. In one case by moving the compute, in the other by holding back the data.
I have already made that bet, on a small scale — a server in the living room, my data on my own disks, and more and more models running on top of it. Not because the cloud is the enemy, but because it seems sensible that at least part of my digital infrastructure shouldn’t depend on somebody else’s price list.
If AI really does become the infrastructure of everything, the question that matters will not be “how smart is it?”.
It will be: whose machine is it running on?
Sources and further reading
Starting point
- Simone Rizzo — I modelli AI sono ormai tutti intelligenti. Ora cambia tutto. (analysis of the open source / hardware / bubble picture, in Italian)
- Decodifica — Perché dovresti scommettere la tua carriera sull’AI locale? (a career perspective on on-premise AI, in Italian)
Models and pricing
- OpenAI cuts GPT-5.6 Luna API prices by 80% — CNBC
- Chinese AI models now capture up to 46% of US enterprise token usage — CNBC investigation, July 2026
- Kimi K3: open weights, 2.8T parameters — Moonshot AI
- GLM-5.2 benchmark deep dive and The Open Weight Models that Matter — OpenRouter
- Inkling: Our Open-Weights Model — Thinking Machines Lab
- Microsoft Build 2026: MAI keynote and the paper Building a Hill Climbing Machine
- Alibaba Qwen3.8-Max reactions: “an API business model wearing an open source jacket” — The New Stack
Industrial policy and markets
- Nvidia, Microsoft, Meta warn against ‘premature restrictions’ of open-weight models — CNBC
- Nvidia and 24 other companies sign open-weights letter — Tom’s Hardware
- The AI capex-to-revenue gap is widening — Forbes
- Tracker of AI company spend vs revenue (estimates, to be read as such)
- Where China’s AI chip supply chain stands in 2026
Hardware and memory
- 2025–present global memory supply shortage — Wikipedia
- Memory price surge begins to cool — Tom’s Hardware
- Apple Upgrade launches in the United States — Apple Newsroom
- NVIDIA DGX Station for Windows — official specifications
- Apple’s rumored M7 Ultra targets 1.5TB of memory — Tom’s Hardware (rumour, unconfirmed)
Local inference and harness
- antirez/ds4 — DwarfStar 4 — open source inference engine in C
- How enabling two settings tripled our scores on the ARC-AGI-3 benchmark — OpenAI
- Self-Hosted LLM Guide 2026
Local pseudonymisation
- Rizzo-AI-Academy/rizzo-pii — project repository (MIT)
- rizzoaiacademy/rizzo-pii-0.3B — model card, metrics and supported categories
- Local, reversible PII anonymisation — project documentation
- Regulation (EU) 2016/679 — Art. 4(5) and Recital 26 — definition of pseudonymisation and the scope of personal data
Regulation, sovereignty and jobs
- Data Sovereignty for Enterprise AI: Complete Guide 2026 — NeuralTrust
- Sovereign AI: The 2026 Enterprise Guide for Regulated Industries
- Data Scientists — Occupational Outlook Handbook and Information Security Analysts — U.S. Bureau of Labor Statistics, 2024–2034 projections