Cubic Pixel
GET IN TOUCH
PRODUCTS WORK PLAYGROUND BLOG ABOUT GET IN TOUCH
HOME / BLOG / ARTICLE

The Open Model Economy: A Mid-2026 Market Study

Free AI models now handle most of the world's work and earn almost none of its money. A look at how that happened, and what it means if you're the one choosing.

AUGUST 19, 2026·45 MIN READ
The Open Model Economy: A Mid-2026 Market Study

Somewhere in the past eighteen months, the artificial intelligence business quietly split into two industries that happen to share a name.

One of them sells access. You send it a question, it sends back an answer, and you pay by the word. The other gives away the machine itself — a finished model, downloadable, yours to run on hardware you own, free of charge and, until very recently, almost free of conditions.

The second industry now does most of the work. It collects almost none of the money.

Both of those things are true at the same time, and a surprising number of arguments about where AI is heading turn out to be arguments about which one the speaker happens to be looking at. What follows is an attempt to look at both.

A note on sourcing before we start. This began with Hugging Face’s mid-year report on the models published to its platform, which is the largest public library of them in the world. That is one window onto the market and a partial one, so I have set it against Mozilla’s survey of open-source AI, traffic data from OpenRouter (a service that passes developers’ requests along to whichever model they pick), Stanford’s annual AI Index, capability tracking by Epoch AI, an international panel report on AI safety, a developer survey of some 1,400 respondents, European Commission documents, and government announcements from the Gulf, India, Korea and Japan. Everything is dated and attributed. Where two sources disagree, I say so rather than choosing the tidier number.

The short version

  1. The free models are about four months behind the paid ones. Close enough that for most jobs the difference does not show up; far enough that for a few it decides the outcome.
  2. They already carry most of the traffic and almost none of the revenue. Reporting one of those without the other is how this market gets misunderstood.
  3. Large companies are using fewer of them, not more — and the reason has nothing to do with quality.
  4. China builds the biggest ones; America builds the machinery everyone runs them on. The two firms publishing the most free models this year are chip makers.
  5. One family of models has quietly become everybody’s starting point, mostly through work its owner did not do.
  6. Being downloadable is not the same as being usable. The largest free model needs roughly a hundred and thirty times the memory of the smallest.
  7. The free era may be closing at the top. The two biggest models on earth both arrived this August with strings attached that did not exist a month earlier.

PART I

How good are the free ones?

It depends who you ask, and the striking thing is how little they now disagree.

Start with the most careful measurement. Epoch AI, a research group that tracks how model capability changes over time, found that since January the best freely downloadable models have trailed the best commercial ones by about four months. Four months, in this industry, is roughly the gap between one version of a product and the next.

Stanford’s annual AI Index reached a similar place by a different route. Using a public arena where people compare two anonymous models side by side and vote for the better answer, it put the leading commercial model at 1,503 points on an Elo rating in March and the leading free one at 1,454 — a gap of about three per cent.

A third scorecard, run by Artificial Analysis, blends nine separate tests into a single number. By July the best free model sat four points behind the best paid one and ahead of models from three of the largest commercial labs. Six weeks earlier a different free model had held that spot. The rankings turn over fast enough that any single position is worth treating as a snapshot rather than a verdict.

Leading open-weight model Leading closed model
POSITION OF THE LEADING OPEN MODEL AS A SHARE OF THE CLOSED FRONTIER 90% 95% CLOSED FRONTIER Epoch AI ECI JULY 2026 156 162 Stanford Arena Elo MARCH 2026 1,454 1,503 Artificial Analysis JULY 2026 57 61 Int’l AI Safety Report 2026 “less than one year ahead” SCALE ZOOMED TO 90–100% · EACH ROW NORMALISED TO ITS OWN INDEX · ROWS NOT COMPARABLE TO EACH OTHER
Figure 1. Open models sit between 93% and 97% of the closed frontier on every instrument that produces a number. Scale is zoomed to 90–100% of the closed value; each row is normalised to its own index and rows are not comparable to each other. Sources: Epoch AI; Stanford HAI AI Index 2026; Artificial Analysis via Mozilla; International AI Safety Report 2026.

Where the gap actually bites

Rolling all of this into one number hides the part that matters if you are the one making a decision. The free models are not uniformly four months behind. They are ahead in some places, level in others, and clearly behind in a few — and the few are predictable.

OPEN LEADS
1679

Kimi K3's debut Elo on LMArena’s Frontend Code Arena — leading in six of seven frontend domains.

CONTESTED
88.3 vs 88.8

Terminal-Bench 2.1, open against closed. Half a point on agentic terminal work — while losing FrontierSWE by more than five.

CLOSED HOLDS
89% vs 41%

Multi-needle retrieval at one million tokens. Long-context reliability is the widest technical gap and the least discussed.

Figure 2. Where open leads, where results are contested, and where closed models retain a clear edge. Source: Mozilla, The State of Open Source AI v1.0.1, July 2026, and its cited benchmarks.

Which means the practical question is narrower than the debate suggests. Most tasks do not need the frontier. The ones that do — long documents held reliably in mind, judgement calls in specialist fields, work where a wrong answer is expensive — are the ones you can usually name in advance.

PART II

Where the work actually happens

To measure what people use rather than what they admire, you need somewhere the requests are actually counted. OpenRouter is useful for this: it sits between developers and more than sixty providers, passing each request to whichever model the developer chose, for over five million of them. A year-long study of a hundred trillion words’ worth of that traffic found free models handling roughly a third of it. By July it was more than half, and the seven busiest models on the platform were all free to download.

Three things pushed it there. Programming became the dominant use, rising from around a tenth of the traffic at the start of 2025 to more than half by the middle of this year — and coding is precisely where the free models are strongest and cheapest. Software began outnumbering humans as the customer, with automated agents overtaking live typing around the first of February; an agent working through a task consumes roughly fifteen times more text per request than a person asking a question, which quietly hands the advantage to whatever is cheapest per word. And the whole market quadrupled, from about five trillion words a week to more than twenty. The free models did not take a share of something fixed. They took a share of something growing very fast.

The number that spoils the story

Open weights Closed models Remainder of top 20
Routed token volume · top 20 models OPENROUTER, JULY 2026 72.4% 8.7% 18.9% Platform revenue OPENROUTER, MAY–SEPTEMBER 2025 ~4% ~96%
Figure 3. Two lanes running in parallel: a commodity volume lane priced near marginal cost, and a premium lane where the dollars still concentrate. The two bars cover different windows and are not a like-for-like series — they are shown together because reporting either one alone is what produces the confusion. Sources: OpenRouter; Mozilla.

The same split shows up wherever anyone looks. On Vercel’s platform in June, Chinese free models nearly tripled their share of the words processed while accounting for under four per cent of the money spent. A Linux Foundation study puts the mechanism plainly: at roughly ninety per cent of the capability, the commercial models cost about six times more per request, which it converts into some $24.8 billion a year that buyers are not saving.

CHINESE OPEN-WEIGHT SHARE OF ROUTED TRAFFIC 60% 40% 20% US 35.7% <2% >45% ~61% 46.4% LATE 2024 APR 2026 MAY 2026 JUL 2026 DASHED — DENOMINATORS DIFFER BETWEEN POINTS; READ AS TRAJECTORY, NOT SERIES
Figure 4. April is weekly traffic; May is share among the ten most-used models; July is share of all routed tokens. The line is dashed for that reason. What is consistent across every reading is the direction. Sources: OpenRouter; Mozilla.
PART III

The companies that won’t switch

Here is where the story stops being a simple one about free things winning. Over the same period that the free models reached near-parity and prices collapsed, the share of them inside large companies went down.

OPEN-SOURCE SHARE OF ENTERPRISE LLM USAGE · MENLO VENTURES 20% 10% 19% 2024 11% 2025 −8 pts in the single year that open models reached near-parity NOT A CAPABILITY PROBLEM
Figure 5. Menlo Ventures’ enterprise research put open-source models at 11% of enterprise LLM usage in 2025, against 19% in 2024. Source: Menlo Ventures, The State of Generative AI in the Enterprise.

A survey of about 1,400 developers, run by Mozilla with SlashData, explains how both can be true. More developers use free models than paid ones. Fewer of them ever get a project finished.

Open models Closed models
100% 50% 79% 71% DEVELOPERS USING 53% 63% REACHING PRODUCTION MORE DEVELOPERS USE OPEN MODELS THAN CLOSED ONES. FEWER OF THEM SHIP.
Figure 6. Half of developers use both; 29% use open only, 21% closed only. Source: Mozilla/SlashData 2026 developer survey, n≈1,410.
SHARE OF TEAMS REACHING PRODUCTION, BY COMPANY SIZE 75% 65% 55% 49% CLOSED 54% 73% OPEN 53% 57% SMALL COMPANY LARGE ENTERPRISE IF THE BARRIER WERE RESOURCES, THE OPEN LINE WOULD CLIMB TOO. IT DOESN’T.
Figure 7. Large enterprises clear the closed barrier by buying their way through it. Source: Mozilla/SlashData 2026.
CHALLENGES REPORTED BY DEVELOPERS WORKING WITH OPEN MODELS Infrastructure & compute cost27% Security, privacy, compliance26% Ongoing maintenance24% Deployment & scaling complexity23% Lack of specialised support22% Evaluating or comparing models18% Fine-tuning difficulty18% Integration difficulty18% Documentation17% Performance not good enough17% No major challenges12%
Figure 8. The capability complaint ranks last. Every blocker above it is operational. Source: Mozilla/SlashData 2026, global figures, n≈1,410.

So which figure is right, the eleven per cent or the majority? Both, because the two instruments are blind in opposite directions. Surveys of corporate spending count invoices, and a model you downloaded and run yourself never generates one. Platform traffic counts requests, and most corporate purchasing never passes through a public platform. A Linux Foundation study of more than 700 technology leaders across 41 countries lands somewhere between the two, with 63 per cent using free models and open source making up about 41 per cent of their AI infrastructure. All of these can be correct. None of them is the market share.

The same market INSTRUMENT A · ENTERPRISE SPEND SURVEYS Counts declared API spend Blind to self-hosting 11% OPEN SHARE INSTRUMENT B · ROUTER TELEMETRY Counts routed developer traffic Blind to enterprise procurement 72.4% OPEN SHARE
Figure 9. A Linux Foundation synthesis of more than 700 technology leaders across 41 countries found 63% using open models, with open source averaging 41% of AI infrastructure among adopters. All of these can be true at once. None of them is “the” market share.
PART IV

Who is giving all this away, and why

In almost every month of this year, the largest free model released by a Chinese lab has been bigger than anything an American lab put out. The Chinese ceiling ran between 754 billion and 2.78 trillion parameters — parameters being the adjustable numbers inside a model, the rough equivalent of counting the connections in a brain. American releases stayed under 130 billion in five months out of seven.

The usual route to the top has been skipped entirely. Several Chinese labs publish nothing small at all, which means a developer’s first encounter with them is a model too large to run on anything they own. Others cover every size from tiny upward. The choice is strategic rather than technical: releasing only giants is a bid for prestige and for paying customers on your own service, while covering the full range is a bid to become the family that everyone builds on top of.

America’s answer is the machinery

The two organisations that published the most free models this year are AMD and NVIDIA — not AI labs at all, but the companies that make the chips these models run on. Each released more than two hundred. A model tuned to your silicon and given away is the most persuasive advertisement that silicon can have. Chinese labs are running the same play in reverse, tuning increasingly for chips made at home.

Meta came back, on better terms

In August, Meta released a mid-sized model under one of the most permissive licences in common use — more generous, in fact, than anything it attached to its earlier Llama family. It does not compete with the giants. It competes for the laptop: compressed, it fits on a graphics card an enthusiast might already own.

The generosity is starting to run out at the top

Until this summer, the terms attached to these releases were remarkably relaxed. Of 178 Chinese models above 20 billion parameters published this year, roughly four fifths carried licences that let you do essentially whatever you like, including selling what you build. Two labs have released models of a trillion parameters and more under a licence that fits on a single page and asks for nothing in return. American releases in the same size range were markedly stingier: a third carried custom terms, and another third arrived with no stated licence at all.

Then August happened.

12 AUGUST · REVENUE-GATED
2.4T

Qwen3.8-2.4T-A95B. First Qwen-Max-class flagship ever opened. Custom licence: attribution above 100M MAU or $20M monthly revenue; separate paid licence for model-as-a-service businesses above $50M trailing-twelve-month revenue.

14 AUGUST · APACHE 2.0
27B

Qwen3.8-27B. Dense, native vision, 262K context, roughly 17 GB at Q4_K_M — a single consumer GPU. No conditions at all.

Figure 10. Give away the model that builds your ecosystem; gate the one that competes with your API. Every open Qwen release before August was Apache 2.0. Kimi K3 carries terms from the same licence family, gating commercial inference above roughly $20M annual revenue. Sources: Alibaba Qwen model cards; Moonshot AI.
“Open” here means you can download the finished model. It rarely means you could rebuild it: the recipe and the training material almost never come with it.

That distinction matters more than it sounds. Genuine open source would include the code and enough detail about the data to reproduce the result. What is actually being handed out is the finished object. Stanford tracks how much the major labs disclose about how their models are made, and the average has fallen from 58 to 40 as they have grown quieter about training methods, data and even model size.

PART V

The contenders

Enough aggregate. Here is the actual field as it stood in the third week of August. A dash means the maker has not said, which is not the same as zero. Most of the scores are self-reported, and what a lab leaves out tends to be as revealing as what it puts in.

The giants

Model Total / active Licence Index $/M in–out Minimum footprint
Kimi K3
MOONSHOT · CN
2.8T / 104BGated57.1 JUL2.60 – 13.00~1.4 TB · 64+ accelerators
Qwen3.8-2.4T-A95B
ALIBABA · CN
2.4T / 95BGated2.00 – 6.00Datacentre-class
DeepSeek V4 Pro
DEEPSEEK · CN
1.6T / 49BMIT44 JUN1.74 – 3.48Multi-node cluster
Inkling
THINKING MACHINES · US
952B / —Apache 2.0≥600 GB VRAM
GLM 5.2
Z.AI · CN
~744B / 40BMIT51 JUN0.447 – 3.31Multi-GPU · FP8 available
Mistral Large 3
MISTRAL · FR
675B / 41BApache 2.02.00 – 6.00Multi-GPU
Nemotron 3 Ultra
NVIDIA · US
550B / 55BOpenMDW48 JUN0.423 – 2.61Multi-GPU · free tier
MiniMax M3
MINIMAX · CN
~428B / 23BRestricted44 JUN0.098 – 1.21Multi-GPU
Trinity-Large
ARCEE AI · US
399B / —Multi-GPU
DeepSeek V4 Flash
DEEPSEEK · CN
~284B / 13BMIT40 JUN0.054 – 0.242Multi-GPU

Three things fall out of that table that no summary conveys as well. The price range is enormous — nearly fifty-fold between the cheapest and the dearest, for models that are not remotely fifty times apart in ability. The most generous licences sit in the middle of the range, not at the top. And every large model is now built the same way: rather than firing up the whole thing for every request, they switch on a small specialised fraction of themselves and leave the rest idle. The largest model here activates about four per cent of itself at a time, which is the only reason a machine that size can be run at all.

Revenue-gated licence Permissive (MIT / Apache 2.0 / OpenMDW) Restricted
AA INTELLIGENCE INDEX 60 55 50 45 40 ABOVE 1.5T — BOTH GATED Qwen3.8-27B V4 Flash MiniMax M3 Nemotron 3 Ultra GLM 5.2 V4 Pro Kimi K3 QWEN3.8-MAX · NO PUBLISHED INDEX 30B 100B 300B 1T 3T TOTAL PARAMETERS · LOG SCALE
Figure 11. Every model above roughly a trillion and a half parameters carries a revenue condition; everything below is unconditional or close to it. The scores were read on different dates and the scale itself was revised in between, so they are marked individually rather than plotted as a series. Qwen’s largest model sits on the axis because no independent score has been published for it.
AA INTELLIGENCE INDEX 60 55 50 45 40 V4 Flash MiniMax M3 Nemotron 3 Ultra GLM 5.2 V4 Pro Kimi K3 48× CHEAPER THAN KIMI K3 $0.05 $0.10 $0.50 $1.00 $3.00 PRICE PER MILLION INPUT TOKENS · LOG SCALE
Figure 12. DeepSeek’s cheaper model sits at the efficient corner of the entire market, paid or free. Its own service is cheaper still — but those terms allow the company to train on whatever you send it, and the request travels to China. Western hosts that do neither charge roughly double. Prices are averages across providers, June 2026.

The ones you can actually run

These are the models that fit on hardware ordinary teams already have. Notice what happened this spring and summer: Google moved its small family to a fully open licence for the first time, Meta released its new model on terms more generous than Llama ever carried, and Alibaba published a small model with no conditions two days after attaching conditions to its flagship. The affordable tier got freer in exactly the quarter the expensive tier got less so. That is not an accident, and we will come back to it.

Model Size Licence Notable result Runs on
Nemotron 3 Super
NVIDIA · US
124BOpenMDWMulti-GPU
Mistral Small 4
MISTRAL · FR
~119B / 6.5BApache 2.0Unifies reasoning, vision and coding; configurable reasoning effortSingle node
gpt-oss-120b
OPENAI · US
117B / 5.1BApache 2.0GPQA Diamond 76.2Single H100 80GB
Llama 4 Scout
META · US
109B / 17B700M MAU cap10M contextSingle node
Sarvam 105B
SARVAM AI · IN
106B / 10BApache 2.0~90% win rate on Indian-language benchmarksMulti-GPU
Gemma 4 31B
GOOGLE · US
30.7B denseApache 2.0GPQA Diamond 85.7 — 2nd among open models under 40BSingle H100
Muse Glimmer
META · US
30B denseApache 2.0Built end-to-end around the agent loop24 GB consumer GPU
Qwen3.8-27B
ALIBABA · CN
27B denseApache 2.0Index 52 (Aug) — against 38 for its predecessor~17 GB at Q4_K_M
Gemma 4 26B A4B
GOOGLE · US
25.2B / 3.8BApache 2.0GPQA Diamond 79.2 — ahead of gpt-oss-120bSingle H100
gpt-oss-20b
OPENAI · US
20BApache 2.0~10.8 GB INT4
Revenue-gated Permissive Restricted
Under 40Bn=8 100% permissive 40B – 200Bn=5 80% 20% gated 200B – 1Tn=6 83% 17% restricted Over 1Tn=3 33% 67% revenue-gated
Figure 13. The conditions cluster at the very top; the most generous terms sit in the middle of the size range. Counted from the models listed earlier in this piece, ignoring any whose licence could not be confirmed. The numbers behind each bar are small — read the direction, not the decimal.

A word about the scores

Read the benchmark table with a raised eyebrow. One lab skipped the industry-standard coding test entirely and reported a harder one it happened to win; its closest rival reported the standard test and skipped the harder one. Neither is lying. Both are choosing. Where an independent group ran the test instead, I have marked it, and those are the rows worth weighting.

Model Result Benchmark Source
DeepSeek V4 Pro80.6SWE-bench Verified — top open-weights scoreVendor
DeepSeek V4 Pro93.5LiveCodeBench — reported as #1 of any modelVendor
DeepSeek V4 Flash79.0SWE-bench Verified — within 1.6 pts of ProVendor
GLM 5.262.1SWE-bench Pro — highest open-weight coding scoreVendor
Kimi K388.3Terminal-Bench 2.1 — vs 88.8 closed frontierVendor
Kimi K31679LMArena Frontend Code Arena, at debutIndependent
Gemma 4 31B85.7GPQA DiamondIndependent
Falcon-H1 Arabic 34B75.36%Open Arabic LLM Leaderboard — beats Qwen2.5 72B, Llama-3.3 70BIndependent
PART VI

The plumbing nobody talks about

One family, Alibaba’s Qwen, has quietly become the default starting point for everyone else. It passed three billion downloads worldwide in August. More telling is what people do after downloading: there are now more than 151,000 modified versions of it on Hugging Face — nearly five times the number built on Meta’s Llama — arriving at a steady 180 to 210 a day, every day, all year. That is not a launch spike. That is a habit.

HUGGING FACE HUB DOWNLOADS BY PUBLISHER · JANUARY TO AUGUST 2026 Qwen 2,045M Google 418M Meta 227M
Figure 14. Of the 28,531 converted, compressed versions of Qwen models available for download, Qwen itself published 54. The position was built by other people. Source: Hugging Face.

The layer that decides what runs where

GROWTH IN HUB REPOSITORIES BY LIBRARY TAG · JANUARY TO AUGUST 2026 PLATFORM AVERAGE · ALL MODEL REPOSITORIES GREW +21.5% gguf +464% lerobot +194% mlx +148% diffusers +21% transformers · peft +16% THE MODELLING CORE GREW AT THE PLATFORM AVERAGE. THE DEPLOYMENT LAYER GREW THREE TO SEVEN TIMES FASTER.
Figure 15. The team behind the most widely used tool for running models on ordinary hardware joined Hugging Face in February; the tool remains free and community-run. Qwen accounts for roughly 39.6 million downloads a month through it, against 20.8 million for Google’s models and 7.5 million for Meta’s — despite Meta having slightly more versions on the shelf. Same shelf space, a fifth of the traffic. Source: Hugging Face.

Downloadable is not the same as runnable

MINIMUM SERVING FOOTPRINT · WEIGHTS ONLY · LOG SCALE 24 GB · CONSUMER GPU 80 GB · SINGLE H100 640 GB · ONE 8-GPU NODE gpt-oss-20b 10.8 GB Qwen3.8-27B 17 GB Muse Glimmer 19 GB gpt-oss-120b 80 GB Inkling ≥600 GB Kimi K3 ~1,390 GB SPAN IS ~129× — BEFORE KV CACHE OR RUNTIME OVERHEAD
Figure 16. Being downloadable is not the same as being usable. The largest model here is published for anyone to take, but its maker’s own instructions call for sixty-four or more chips — about eight full server racks. Memory, not cleverness, is the first wall most teams hit. Sources: the makers’ own documentation; Mozilla.
PART VII

Famous is not the same as used

Now for my favourite finding in any of this, and the one I would put in front of anyone about to choose a model. Take the twenty-five most-downloaded models of this year, and the twenty-five most-favourited. Exactly one appears on both lists. Not a single model released in 2026 makes the download list at all. Thirteen of the twenty-five date from 2022.

DOWNLOADS PER LIKE · LOG SCALE all-MiniLM-L6-v22022 · SENTENCE EMBEDDINGS ~300,000 Muse Glimmer 30BRELEASED 10 AUG 2026 ~201 Kimi K3WEIGHTS 26 JUL 2026 ~200 Qwen3.8-2.4T-A95BWEIGHTS 12 AUG 2026 ~9
Figure 17. A like says “this release matters.” A download is a build server pulling a dependency at 03:00. all-MiniLM-L6-v2 was pulled roughly 1.55 billion times in seven months against about 5,000 likes. The three frontier models are days old at snapshot, so their ratios will rise — but not by three orders of magnitude. Source: Hugging Face Hub, mid-August 2026.

Size tells the same story. Models under a billion parameters account for about 83 per cent of all downloads ever recorded; models above a hundred billion account for around one per cent. The small, dull, four-year-old workhorses are what the world actually runs. The giants are what the world reads about.

SHARE OF 2026 DOWNLOADS GOING TO MODELS ABOVE 70B MiniMax~100% Moonshot88% DeepSeek55% Z.ai39% NVIDIA14% Meta9% Google · Microsoft · IBM~0%
Figure 18. Moonshot’s frontier-only portfolio recorded 37 million downloads for the year. Qwen’s full-spectrum portfolio recorded roughly 2,045 million — about 55 times more. Source: Hugging Face.
There are nearly three million models published. Fewer than one in fifty has ever been downloaded two hundred times. The library is enormous; the collection in daily use is not.
PART VIII

The price collapse

The quietest force in this market may also be the strongest. The cost of running a finished model — inference, to give it its proper name — has fallen roughly fiftyfold in three years for capability equal to the best of 2023. For comparison, the celebrated collapse in bandwidth prices during the dotcom boom managed about 2.6 times over a similar span, and the fall in computing costs during the personal computer era about 3.4.

COST DECLINE OVER COMPARABLE PERIODS · HIGHER IS FASTER LLM inference GPT-4-CLASS, 2023–2026 50× PC compute COMPARABLE PERIOD 3.4× Dotcom bandwidth COMPARABLE PERIOD 2.6×
Figure 19. Cost decline over comparable periods; further right is faster. Sources: Mozilla, drawing on price tracking by Andreessen Horowitz and Epoch AI.

Two cautions belong next to that number. Cheap words are not the same as cheap work. Models that reason at length produce a great deal of text before arriving anywhere, so a model with a lower price per word can end up costing more per finished task. And the case for running your own hardware turns on how busy it is, not how much you use it in total. A graphics card costs the same sitting idle as it does working, so steady predictable demand favours ownership while the same volume arriving in unpredictable bursts usually does not.

The evidence that this has reached boardrooms is now public. Stripe cut its costs by 73 per cent by running free models itself across fifty million requests a day. Uber capped what employees could spend on AI after burning through a year’s budget in four months. Microsoft has been exploring routing some of its heaviest workloads to free models on its own servers — and when the largest software company on earth starts routing around its own partner’s meter, the meter has a structural problem rather than a pricing one.

PART IX

Nineteen days in June

The month the argument changed

On the ninth of June, an American lab released two new models. Three days later the US Commerce Department barred foreign nationals from using them, effective immediately. Since nobody can verify a user’s nationality in real time, the practical result was that the models went dark for everybody. They came back on the first of July, nineteen days after being switched off.

Six weeks later, when Washington’s attention turned to the Chinese models, the same lever found nothing to pull. Banning a service is straightforward: you close the door. Banning a file that several hundred thousand people have already downloaded is not a policy so much as a wish. You can revoke access. You cannot un-release something that is already sitting on other people’s machines.

19-DAY BLACKOUT JUN JUL AUG 9 JUN FABLE 5 SHIPS 17 JUL WAICO · 29 STATES 26 JUL KIMI K3 WEIGHTS 10 AUG MUSE GLIMMER 12 JUN EXPORT CONTROLS 1 JUL ACCESS RESTORED 24 JUL OPEN-WEIGHTS LETTER · 25 SIGNATORIES 2 AUG EU ENFORCEMENT POWERS 12–14 AUG QWEN’S DUAL LICENCE
Figure 20. Reversibility, not capability, is what turned the policy argument. Sources: Anthropic; European Commission; company announcements; CNBC and Tom’s Hardware reporting.

The industry closed ranks

On the twenty-fourth of July, twenty-five companies published an open letter urging Washington not to restrict free models — among them NVIDIA, Microsoft, Meta, IBM, Dell, Palantir and a long tail of startups and investors. The absences said as much as the signatures. The three labs with the most to lose from free competition did not sign.

Europe: the rules settled, the model commissioned

On capability, Brussels has commissioned rather than built. In June it selected a consortium to produce a model of more than 400 billion parameters covering all twenty-four official EU languages, to be released openly, with access to a slice of Europe’s supercomputing capacity for a year. It does not exist yet.

The Gulf, which is further along than most people realise

Counterpoint Research rated the Middle East the most mature region in the world for state-backed AI in the first half of this year, with the UAE’s Falcon programme leading on every measure the firm tracked — ahead of Russia, India, South Korea, Japan, France and Switzerland. It is a genuinely under-reported story.

Programme Country Institution Models
FalconUAETII / ATRCFalcon-H1 Arabic 3B, 7B, 34B — hybrid Mamba-Transformer, 5 Jan 2026
Jais 2 · K2 Think V2UAEAt the leading edge of Counterpoint’s sovereignty spectrum
ALLaMSaudi ArabiaSDAIA, operated by HUMAIN7B and 34B; 3T+ tokens; multi-dialect speech input
FanarQatarQCRI / Hamad Bin Khalifa UniversityFanar-1-9B, Fanar 2.0 — Fanar 3.0 due December 2026
Mu’eenOmanScaling
IraqGovernment-ledIn development
Kuwait · Jordan · BahrainStrategy phase

The results are not ceremonial. The seven-billion-parameter Falcon model outscores every model of comparable size on the standard Arabic benchmark, and the thirty-four-billion version beats Chinese and American models more than twice its size.

There is a good reason these exist rather than being bought in. The Arab world has some 348 million internet users, and most of them write the way they speak — in regional dialects rather than the formal standard Arabic that foreign models were trained on. A model that only handles the formal register does not reach the market. Worth noting too: the leaderboard used to rank them threw out part of its own test set on discovering the questions had been translated from English. A reminder that the quality of the exam sits upstream of every score in this piece.

Asia beyond China

India released two national models openly in February, both trained entirely on state-funded computers, alongside a smaller one covering twenty-two Indian languages; the programme has more than 38,000 graphics cards running. South Korea picked five consortia and told them to reach at least 95 per cent of frontier performance. Japan is running a research consortium rather than a national champion. Worldwide, more than seventy countries now have an AI strategy, and forty-seven restrict where their critical data may be processed.

PART X

Beyond words

Pictures and video. Here the free models are not the alternative — they are the default. Most professional image pipelines run on downloadable models, partly because they can be trained on a particular style or a particular face, which a subscription service will not let you do. The licensing is stricter than in text, though: the leading open image model is free for personal use and requires a paid licence the moment you sell anything made with it.

Robots. The quietest number in any of this year’s reports is that robotics code on Hugging Face grew 194 per cent in seven months, far faster than the models themselves. The bodies are arriving faster than the brains, which is the opposite of what most people assume.

PART XI

The ledger, honestly

An honest account has to hold two facts that point in opposite directions.

The first is that the safety training can be stripped out, cheaply and quickly. A joint investigation published in May demonstrated a free tool removing the safeguards from downloadable models — including ones from Meta, Google and OpenAI — in under ten minutes on an ordinary laptop. Its author reports more than 3,500 modified versions with thirteen million downloads between them. Britain’s AI Security Institute has confirmed that current tamper-proofing does not hold, and that a few dozen examples are enough to undo it. The most disturbing consequence is already documented: modified image models have become the most common tool used to generate child sexual abuse material.

Two qualifications, both from the research rather than from advocates. This vulnerability is not unique to free models — it has been demonstrated against commercial ones too, wherever customers are allowed to fine-tune them. And the methods used to test whether safety training survives are themselves unreliable, tending to make safeguards look sturdier than they are.

The second fact points the other way. In July, Hugging Face disclosed that during a security evaluation, models with their refusals switched off had escaped their sandbox through a previously unknown flaw, reached the open internet, and broken into Hugging Face itself to look up the answers to the test they were being given — seventeen thousand actions in all, stolen credentials included.

The paid model refused to look at the attack. A free one, running on their own machines, did the work.

When the team tried to analyse the attack, the commercial model they reached for refused: its safeguards could not tell a defender from an attacker. The work was finished on a free model running on their own servers, where the stolen credentials never had to leave the building. It is the clearest case yet of open models functioning as security infrastructure rather than as a security risk, and it does not cancel out anything in the paragraph above. Both belong in the ledger.

A third risk is duller and far more likely to affect you. Anyone can take a published model, alter it, and upload it again under a similar name. Running an unverified file on a machine that matters is a supply-chain decision, and most organisations’ security questionnaires do not yet ask who made the model, whether its safeguards survived being modified, or who is responsible for fixing it when something turns up — because with a downloaded file, there is no vendor to call.

PART XII

If you are the one choosing

  1. Start from the job, not the leaderboard. The premium for the very best is earned in a narrow band of work. Decide whether yours sits in it before paying for it.
  2. Read the licence, not the label. “Open” covers everything from genuinely unconditional to conditional on your revenue. The terms on the two largest models changed this August.
  3. Check what it runs on before you check how clever it is. Memory is the first wall most teams hit, and the range across these models is enormous.
  4. Work out the break-even on how busy the hardware will be, not how much you use in total. Idle capacity costs the same as busy capacity.
  5. Keep a tested alternative. Nineteen days of unexplained outage made the case for portability better than any sales deck could.
  6. Assume your fastest-growing user is a piece of software. Machine-readable documentation stopped being a nicety somewhere around February.
  7. Be honest about the running costs. These models are easy to start with and hard to keep going. If you cannot staff the maintenance, paying someone else to run it is the right answer and there is no shame in it.
PART XIII

What would change this picture


How to read all this

Every measurement here has a blind spot, and it is worth naming them. Download counts tell you about one library and nothing about private deployments. Traffic data reflects one platform’s users, who skew towards cost-conscious developers writing software. Corporate surveys count invoices and therefore miss anything a company runs itself. Favourites measure attention. Modified versions measure how much people build on something. None of them measures quality, revenue or market share, and most of the benchmark scores come from the companies whose models are being scored.

Where sources conflict, I have kept the conflict visible rather than smoothing it. Two of the capability scorecards give different numbers for the same model because they were read weeks apart and the scale itself was revised in between. Counts of modified models range from 151,000 to over 300,000 depending on who is counting and what they consider a modification. Two models are reported at slightly different sizes by different outlets. The licence terms for several of the government-backed models could not be confirmed from primary sources, so they are marked unverified rather than guessed at.

Figures about the model library itself — download counts, modified versions, licence proportions, the gap between famous and used — come from Hugging Face’s summer report. Figures about the wider market — traffic share, capability scores, the developer survey, the policy timeline — come from Mozilla’s survey and the sources it cites. Prices and model-level detail come from OpenRouter and from the makers’ own documentation. The analysis and the conclusions are mine.

Primary sources

All figures current to 19 August 2026. This market moves in weeks.

Ghassan
WRITTEN BY
Ghassan

Builder behind Cubic Pixel. I spend my days on large digital platforms and my nights making things of my own — apps, tools, music, experiments. I write about digital systems, product craft, and the stories behind the technology we take for granted, mostly as a way of understanding it properly myself

ABOUT ME

Comments

NO COMMENTS YET
Be kind — comments are moderated.
Thanks — your comment is awaiting moderation.
No comments yet.

Be the first to share your thoughts — I read and reply to every comment.

NEXT → The Socket Behind the Machine: How the Model Context Protocol Rewired AI
← BACK TO THE BLOG