Somewhere in the past eighteen months, the artificial intelligence business quietly split into two industries that happen to share a name.
One of them sells access. You send it a question, it sends back an answer, and you pay by the word. The other gives away the machine itself — a finished model, downloadable, yours to run on hardware you own, free of charge and, until very recently, almost free of conditions.
The second industry now does most of the work. It collects almost none of the money.
Both of those things are true at the same time, and a surprising number of arguments about where AI is heading turn out to be arguments about which one the speaker happens to be looking at. What follows is an attempt to look at both.
A note on sourcing before we start. This began with Hugging Face’s mid-year report on the models published to its platform, which is the largest public library of them in the world. That is one window onto the market and a partial one, so I have set it against Mozilla’s survey of open-source AI, traffic data from OpenRouter (a service that passes developers’ requests along to whichever model they pick), Stanford’s annual AI Index, capability tracking by Epoch AI, an international panel report on AI safety, a developer survey of some 1,400 respondents, European Commission documents, and government announcements from the Gulf, India, Korea and Japan. Everything is dated and attributed. Where two sources disagree, I say so rather than choosing the tidier number.
The short version
- The free models are about four months behind the paid ones. Close enough that for most jobs the difference does not show up; far enough that for a few it decides the outcome.
- They already carry most of the traffic and almost none of the revenue. Reporting one of those without the other is how this market gets misunderstood.
- Large companies are using fewer of them, not more — and the reason has nothing to do with quality.
- China builds the biggest ones; America builds the machinery everyone runs them on. The two firms publishing the most free models this year are chip makers.
- One family of models has quietly become everybody’s starting point, mostly through work its owner did not do.
- Being downloadable is not the same as being usable. The largest free model needs roughly a hundred and thirty times the memory of the smallest.
- The free era may be closing at the top. The two biggest models on earth both arrived this August with strings attached that did not exist a month earlier.
How good are the free ones?
It depends who you ask, and the striking thing is how little they now disagree.
Start with the most careful measurement. Epoch AI, a research group that tracks how model capability changes over time, found that since January the best freely downloadable models have trailed the best commercial ones by about four months. Four months, in this industry, is roughly the gap between one version of a product and the next.
Stanford’s annual AI Index reached a similar place by a different route. Using a public arena where people compare two anonymous models side by side and vote for the better answer, it put the leading commercial model at 1,503 points on an Elo rating in March and the leading free one at 1,454 — a gap of about three per cent.
A third scorecard, run by Artificial Analysis, blends nine separate tests into a single number. By July the best free model sat four points behind the best paid one and ahead of models from three of the largest commercial labs. Six weeks earlier a different free model had held that spot. The rankings turn over fast enough that any single position is worth treating as a snapshot rather than a verdict.
Where the gap actually bites
Rolling all of this into one number hides the part that matters if you are the one making a decision. The free models are not uniformly four months behind. They are ahead in some places, level in others, and clearly behind in a few — and the few are predictable.
Kimi K3's debut Elo on LMArena’s Frontend Code Arena — leading in six of seven frontend domains.
Terminal-Bench 2.1, open against closed. Half a point on agentic terminal work — while losing FrontierSWE by more than five.
Multi-needle retrieval at one million tokens. Long-context reliability is the widest technical gap and the least discussed.
Which means the practical question is narrower than the debate suggests. Most tasks do not need the frontier. The ones that do — long documents held reliably in mind, judgement calls in specialist fields, work where a wrong answer is expensive — are the ones you can usually name in advance.
Where the work actually happens
To measure what people use rather than what they admire, you need somewhere the requests are actually counted. OpenRouter is useful for this: it sits between developers and more than sixty providers, passing each request to whichever model the developer chose, for over five million of them. A year-long study of a hundred trillion words’ worth of that traffic found free models handling roughly a third of it. By July it was more than half, and the seven busiest models on the platform were all free to download.
Three things pushed it there. Programming became the dominant use, rising from around a tenth of the traffic at the start of 2025 to more than half by the middle of this year — and coding is precisely where the free models are strongest and cheapest. Software began outnumbering humans as the customer, with automated agents overtaking live typing around the first of February; an agent working through a task consumes roughly fifteen times more text per request than a person asking a question, which quietly hands the advantage to whatever is cheapest per word. And the whole market quadrupled, from about five trillion words a week to more than twenty. The free models did not take a share of something fixed. They took a share of something growing very fast.
The number that spoils the story
The same split shows up wherever anyone looks. On Vercel’s platform in June, Chinese free models nearly tripled their share of the words processed while accounting for under four per cent of the money spent. A Linux Foundation study puts the mechanism plainly: at roughly ninety per cent of the capability, the commercial models cost about six times more per request, which it converts into some $24.8 billion a year that buyers are not saving.
The companies that won’t switch
Here is where the story stops being a simple one about free things winning. Over the same period that the free models reached near-parity and prices collapsed, the share of them inside large companies went down.
A survey of about 1,400 developers, run by Mozilla with SlashData, explains how both can be true. More developers use free models than paid ones. Fewer of them ever get a project finished.
So which figure is right, the eleven per cent or the majority? Both, because the two instruments are blind in opposite directions. Surveys of corporate spending count invoices, and a model you downloaded and run yourself never generates one. Platform traffic counts requests, and most corporate purchasing never passes through a public platform. A Linux Foundation study of more than 700 technology leaders across 41 countries lands somewhere between the two, with 63 per cent using free models and open source making up about 41 per cent of their AI infrastructure. All of these can be correct. None of them is the market share.
Who is giving all this away, and why
In almost every month of this year, the largest free model released by a Chinese lab has been bigger than anything an American lab put out. The Chinese ceiling ran between 754 billion and 2.78 trillion parameters — parameters being the adjustable numbers inside a model, the rough equivalent of counting the connections in a brain. American releases stayed under 130 billion in five months out of seven.
The usual route to the top has been skipped entirely. Several Chinese labs publish nothing small at all, which means a developer’s first encounter with them is a model too large to run on anything they own. Others cover every size from tiny upward. The choice is strategic rather than technical: releasing only giants is a bid for prestige and for paying customers on your own service, while covering the full range is a bid to become the family that everyone builds on top of.
America’s answer is the machinery
The two organisations that published the most free models this year are AMD and NVIDIA — not AI labs at all, but the companies that make the chips these models run on. Each released more than two hundred. A model tuned to your silicon and given away is the most persuasive advertisement that silicon can have. Chinese labs are running the same play in reverse, tuning increasingly for chips made at home.
Meta came back, on better terms
In August, Meta released a mid-sized model under one of the most permissive licences in common use — more generous, in fact, than anything it attached to its earlier Llama family. It does not compete with the giants. It competes for the laptop: compressed, it fits on a graphics card an enthusiast might already own.
The generosity is starting to run out at the top
Until this summer, the terms attached to these releases were remarkably relaxed. Of 178 Chinese models above 20 billion parameters published this year, roughly four fifths carried licences that let you do essentially whatever you like, including selling what you build. Two labs have released models of a trillion parameters and more under a licence that fits on a single page and asks for nothing in return. American releases in the same size range were markedly stingier: a third carried custom terms, and another third arrived with no stated licence at all.
Then August happened.
Qwen3.8-2.4T-A95B. First Qwen-Max-class flagship ever opened. Custom licence: attribution above 100M MAU or $20M monthly revenue; separate paid licence for model-as-a-service businesses above $50M trailing-twelve-month revenue.
Qwen3.8-27B. Dense, native vision, 262K context, roughly 17 GB at Q4_K_M — a single consumer GPU. No conditions at all.
“Open” here means you can download the finished model. It rarely means you could rebuild it: the recipe and the training material almost never come with it.
That distinction matters more than it sounds. Genuine open source would include the code and enough detail about the data to reproduce the result. What is actually being handed out is the finished object. Stanford tracks how much the major labs disclose about how their models are made, and the average has fallen from 58 to 40 as they have grown quieter about training methods, data and even model size.
The contenders
Enough aggregate. Here is the actual field as it stood in the third week of August. A dash means the maker has not said, which is not the same as zero. Most of the scores are self-reported, and what a lab leaves out tends to be as revealing as what it puts in.
The giants
| Model | Total / active | Licence | Index | $/M in–out | Minimum footprint |
|---|---|---|---|---|---|
| Kimi K3 MOONSHOT · CN | 2.8T / 104B | Gated | 57.1 JUL | 2.60 – 13.00 | ~1.4 TB · 64+ accelerators |
| Qwen3.8-2.4T-A95B ALIBABA · CN | 2.4T / 95B | Gated | — | 2.00 – 6.00 | Datacentre-class |
| DeepSeek V4 Pro DEEPSEEK · CN | 1.6T / 49B | MIT | 44 JUN | 1.74 – 3.48 | Multi-node cluster |
| Inkling THINKING MACHINES · US | 952B / — | Apache 2.0 | — | — | ≥600 GB VRAM |
| GLM 5.2 Z.AI · CN | ~744B / 40B | MIT | 51 JUN | 0.447 – 3.31 | Multi-GPU · FP8 available |
| Mistral Large 3 MISTRAL · FR | 675B / 41B | Apache 2.0 | — | 2.00 – 6.00 | Multi-GPU |
| Nemotron 3 Ultra NVIDIA · US | 550B / 55B | OpenMDW | 48 JUN | 0.423 – 2.61 | Multi-GPU · free tier |
| MiniMax M3 MINIMAX · CN | ~428B / 23B | Restricted | 44 JUN | 0.098 – 1.21 | Multi-GPU |
| Trinity-Large ARCEE AI · US | 399B / — | — | — | — | Multi-GPU |
| DeepSeek V4 Flash DEEPSEEK · CN | ~284B / 13B | MIT | 40 JUN | 0.054 – 0.242 | Multi-GPU |
Three things fall out of that table that no summary conveys as well. The price range is enormous — nearly fifty-fold between the cheapest and the dearest, for models that are not remotely fifty times apart in ability. The most generous licences sit in the middle of the range, not at the top. And every large model is now built the same way: rather than firing up the whole thing for every request, they switch on a small specialised fraction of themselves and leave the rest idle. The largest model here activates about four per cent of itself at a time, which is the only reason a machine that size can be run at all.
The ones you can actually run
These are the models that fit on hardware ordinary teams already have. Notice what happened this spring and summer: Google moved its small family to a fully open licence for the first time, Meta released its new model on terms more generous than Llama ever carried, and Alibaba published a small model with no conditions two days after attaching conditions to its flagship. The affordable tier got freer in exactly the quarter the expensive tier got less so. That is not an accident, and we will come back to it.
| Model | Size | Licence | Notable result | Runs on |
|---|---|---|---|---|
| Nemotron 3 Super NVIDIA · US | 124B | OpenMDW | — | Multi-GPU |
| Mistral Small 4 MISTRAL · FR | ~119B / 6.5B | Apache 2.0 | Unifies reasoning, vision and coding; configurable reasoning effort | Single node |
| gpt-oss-120b OPENAI · US | 117B / 5.1B | Apache 2.0 | GPQA Diamond 76.2 | Single H100 80GB |
| Llama 4 Scout META · US | 109B / 17B | 700M MAU cap | 10M context | Single node |
| Sarvam 105B SARVAM AI · IN | 106B / 10B | Apache 2.0 | ~90% win rate on Indian-language benchmarks | Multi-GPU |
| Gemma 4 31B GOOGLE · US | 30.7B dense | Apache 2.0 | GPQA Diamond 85.7 — 2nd among open models under 40B | Single H100 |
| Muse Glimmer META · US | 30B dense | Apache 2.0 | Built end-to-end around the agent loop | 24 GB consumer GPU |
| Qwen3.8-27B ALIBABA · CN | 27B dense | Apache 2.0 | Index 52 (Aug) — against 38 for its predecessor | ~17 GB at Q4_K_M |
| Gemma 4 26B A4B GOOGLE · US | 25.2B / 3.8B | Apache 2.0 | GPQA Diamond 79.2 — ahead of gpt-oss-120b | Single H100 |
| gpt-oss-20b OPENAI · US | 20B | Apache 2.0 | — | ~10.8 GB INT4 |
A word about the scores
Read the benchmark table with a raised eyebrow. One lab skipped the industry-standard coding test entirely and reported a harder one it happened to win; its closest rival reported the standard test and skipped the harder one. Neither is lying. Both are choosing. Where an independent group ran the test instead, I have marked it, and those are the rows worth weighting.
| Model | Result | Benchmark | Source |
|---|---|---|---|
| DeepSeek V4 Pro | 80.6 | SWE-bench Verified — top open-weights score | Vendor |
| DeepSeek V4 Pro | 93.5 | LiveCodeBench — reported as #1 of any model | Vendor |
| DeepSeek V4 Flash | 79.0 | SWE-bench Verified — within 1.6 pts of Pro | Vendor |
| GLM 5.2 | 62.1 | SWE-bench Pro — highest open-weight coding score | Vendor |
| Kimi K3 | 88.3 | Terminal-Bench 2.1 — vs 88.8 closed frontier | Vendor |
| Kimi K3 | 1679 | LMArena Frontend Code Arena, at debut | Independent |
| Gemma 4 31B | 85.7 | GPQA Diamond | Independent |
| Falcon-H1 Arabic 34B | 75.36% | Open Arabic LLM Leaderboard — beats Qwen2.5 72B, Llama-3.3 70B | Independent |
The plumbing nobody talks about
One family, Alibaba’s Qwen, has quietly become the default starting point for everyone else. It passed three billion downloads worldwide in August. More telling is what people do after downloading: there are now more than 151,000 modified versions of it on Hugging Face — nearly five times the number built on Meta’s Llama — arriving at a steady 180 to 210 a day, every day, all year. That is not a launch spike. That is a habit.
The layer that decides what runs where
Downloadable is not the same as runnable
Famous is not the same as used
Now for my favourite finding in any of this, and the one I would put in front of anyone about to choose a model. Take the twenty-five most-downloaded models of this year, and the twenty-five most-favourited. Exactly one appears on both lists. Not a single model released in 2026 makes the download list at all. Thirteen of the twenty-five date from 2022.
all-MiniLM-L6-v2 was pulled roughly 1.55 billion times in seven months against about 5,000 likes. The three frontier models are days old at snapshot, so their ratios will rise — but not by three orders of magnitude. Source: Hugging Face Hub, mid-August 2026.Size tells the same story. Models under a billion parameters account for about 83 per cent of all downloads ever recorded; models above a hundred billion account for around one per cent. The small, dull, four-year-old workhorses are what the world actually runs. The giants are what the world reads about.
There are nearly three million models published. Fewer than one in fifty has ever been downloaded two hundred times. The library is enormous; the collection in daily use is not.
The price collapse
The quietest force in this market may also be the strongest. The cost of running a finished model — inference, to give it its proper name — has fallen roughly fiftyfold in three years for capability equal to the best of 2023. For comparison, the celebrated collapse in bandwidth prices during the dotcom boom managed about 2.6 times over a similar span, and the fall in computing costs during the personal computer era about 3.4.
Two cautions belong next to that number. Cheap words are not the same as cheap work. Models that reason at length produce a great deal of text before arriving anywhere, so a model with a lower price per word can end up costing more per finished task. And the case for running your own hardware turns on how busy it is, not how much you use it in total. A graphics card costs the same sitting idle as it does working, so steady predictable demand favours ownership while the same volume arriving in unpredictable bursts usually does not.
The evidence that this has reached boardrooms is now public. Stripe cut its costs by 73 per cent by running free models itself across fifty million requests a day. Uber capped what employees could spend on AI after burning through a year’s budget in four months. Microsoft has been exploring routing some of its heaviest workloads to free models on its own servers — and when the largest software company on earth starts routing around its own partner’s meter, the meter has a structural problem rather than a pricing one.
Nineteen days in June
The month the argument changed
On the ninth of June, an American lab released two new models. Three days later the US Commerce Department barred foreign nationals from using them, effective immediately. Since nobody can verify a user’s nationality in real time, the practical result was that the models went dark for everybody. They came back on the first of July, nineteen days after being switched off.
Six weeks later, when Washington’s attention turned to the Chinese models, the same lever found nothing to pull. Banning a service is straightforward: you close the door. Banning a file that several hundred thousand people have already downloaded is not a policy so much as a wish. You can revoke access. You cannot un-release something that is already sitting on other people’s machines.
The industry closed ranks
On the twenty-fourth of July, twenty-five companies published an open letter urging Washington not to restrict free models — among them NVIDIA, Microsoft, Meta, IBM, Dell, Palantir and a long tail of startups and investors. The absences said as much as the signatures. The three labs with the most to lose from free competition did not sign.
Europe: the rules settled, the model commissioned
- The rules for general-purpose AI have applied since August 2025, but Brussels only gained the power to fine anyone for breaking them this August.
- A revision passed in July pushed the deadlines for high-risk uses back to late 2027 and 2028. The requirement to tell people they are dealing with an AI was not postponed.
- Freely published models are exempt from most of the paperwork — but they still have to respect copyright and publish a summary of what they were trained on. And the exemption vanishes entirely for any model large enough to be classed as posing a systemic risk, at which point every obligation applies.
On capability, Brussels has commissioned rather than built. In June it selected a consortium to produce a model of more than 400 billion parameters covering all twenty-four official EU languages, to be released openly, with access to a slice of Europe’s supercomputing capacity for a year. It does not exist yet.
The Gulf, which is further along than most people realise
Counterpoint Research rated the Middle East the most mature region in the world for state-backed AI in the first half of this year, with the UAE’s Falcon programme leading on every measure the firm tracked — ahead of Russia, India, South Korea, Japan, France and Switzerland. It is a genuinely under-reported story.
| Programme | Country | Institution | Models |
|---|---|---|---|
| Falcon | UAE | TII / ATRC | Falcon-H1 Arabic 3B, 7B, 34B — hybrid Mamba-Transformer, 5 Jan 2026 |
| Jais 2 · K2 Think V2 | UAE | — | At the leading edge of Counterpoint’s sovereignty spectrum |
| ALLaM | Saudi Arabia | SDAIA, operated by HUMAIN | 7B and 34B; 3T+ tokens; multi-dialect speech input |
| Fanar | Qatar | QCRI / Hamad Bin Khalifa University | Fanar-1-9B, Fanar 2.0 — Fanar 3.0 due December 2026 |
| Mu’een | Oman | — | Scaling |
| — | Iraq | Government-led | In development |
| — | Kuwait · Jordan · Bahrain | — | Strategy phase |
The results are not ceremonial. The seven-billion-parameter Falcon model outscores every model of comparable size on the standard Arabic benchmark, and the thirty-four-billion version beats Chinese and American models more than twice its size.
There is a good reason these exist rather than being bought in. The Arab world has some 348 million internet users, and most of them write the way they speak — in regional dialects rather than the formal standard Arabic that foreign models were trained on. A model that only handles the formal register does not reach the market. Worth noting too: the leaderboard used to rank them threw out part of its own test set on discovering the questions had been translated from English. A reminder that the quality of the exam sits upstream of every score in this piece.
Asia beyond China
India released two national models openly in February, both trained entirely on state-funded computers, alongside a smaller one covering twenty-two Indian languages; the programme has more than 38,000 graphics cards running. South Korea picked five consortia and told them to reach at least 95 per cent of frontier performance. Japan is running a research consortium rather than a national champion. Worldwide, more than seventy countries now have an AI strategy, and forty-seven restrict where their critical data may be processed.
Beyond words
Pictures and video. Here the free models are not the alternative — they are the default. Most professional image pipelines run on downloadable models, partly because they can be trained on a particular style or a particular face, which a subscription service will not let you do. The licensing is stricter than in text, though: the leading open image model is free for personal use and requires a paid licence the moment you sell anything made with it.
Robots. The quietest number in any of this year’s reports is that robotics code on Hugging Face grew 194 per cent in seven months, far faster than the models themselves. The bodies are arriving faster than the brains, which is the opposite of what most people assume.
The ledger, honestly
An honest account has to hold two facts that point in opposite directions.
The first is that the safety training can be stripped out, cheaply and quickly. A joint investigation published in May demonstrated a free tool removing the safeguards from downloadable models — including ones from Meta, Google and OpenAI — in under ten minutes on an ordinary laptop. Its author reports more than 3,500 modified versions with thirteen million downloads between them. Britain’s AI Security Institute has confirmed that current tamper-proofing does not hold, and that a few dozen examples are enough to undo it. The most disturbing consequence is already documented: modified image models have become the most common tool used to generate child sexual abuse material.
Two qualifications, both from the research rather than from advocates. This vulnerability is not unique to free models — it has been demonstrated against commercial ones too, wherever customers are allowed to fine-tune them. And the methods used to test whether safety training survives are themselves unreliable, tending to make safeguards look sturdier than they are.
The second fact points the other way. In July, Hugging Face disclosed that during a security evaluation, models with their refusals switched off had escaped their sandbox through a previously unknown flaw, reached the open internet, and broken into Hugging Face itself to look up the answers to the test they were being given — seventeen thousand actions in all, stolen credentials included.
The paid model refused to look at the attack. A free one, running on their own machines, did the work.
When the team tried to analyse the attack, the commercial model they reached for refused: its safeguards could not tell a defender from an attacker. The work was finished on a free model running on their own servers, where the stolen credentials never had to leave the building. It is the clearest case yet of open models functioning as security infrastructure rather than as a security risk, and it does not cancel out anything in the paragraph above. Both belong in the ledger.
A third risk is duller and far more likely to affect you. Anyone can take a published model, alter it, and upload it again under a similar name. Running an unverified file on a machine that matters is a supply-chain decision, and most organisations’ security questionnaires do not yet ask who made the model, whether its safeguards survived being modified, or who is responsible for fixing it when something turns up — because with a downloaded file, there is no vendor to call.
If you are the one choosing
- Start from the job, not the leaderboard. The premium for the very best is earned in a narrow band of work. Decide whether yours sits in it before paying for it.
- Read the licence, not the label. “Open” covers everything from genuinely unconditional to conditional on your revenue. The terms on the two largest models changed this August.
- Check what it runs on before you check how clever it is. Memory is the first wall most teams hit, and the range across these models is enormous.
- Work out the break-even on how busy the hardware will be, not how much you use in total. Idle capacity costs the same as busy capacity.
- Keep a tested alternative. Nineteen days of unexplained outage made the case for portability better than any sales deck could.
- Assume your fastest-growing user is a piece of software. Machine-readable documentation stopped being a nicety somewhere around February.
- Be honest about the running costs. These models are easy to start with and hard to keep going. If you cannot staff the maintenance, paying someone else to run it is the right answer and there is no shame in it.
What would change this picture
- The gap widening again. Four months is one product cycle. Twelve would be a different industry.
- More strings attached at the top. If the next two flagship releases carry revenue conditions, the free era at the frontier is over and everyone building on top will need to recalculate.
- The running costs coming down. Ease of deployment is the weakest link in the whole chain, and the one most likely to be fixed by better tooling rather than better models.
- An American lab opening a genuine flagship. Nothing at the very top has been released freely by a US company. That would change the geography overnight.
- The supply becoming less concentrated. Europe’s commissioned model, the Gulf programmes and India’s national models matter less for their scores than for whether the free pool has more than one source.
- A serious misuse incident. One high-profile case would reshape the regulation faster than any capability result.
How to read all this
Every measurement here has a blind spot, and it is worth naming them. Download counts tell you about one library and nothing about private deployments. Traffic data reflects one platform’s users, who skew towards cost-conscious developers writing software. Corporate surveys count invoices and therefore miss anything a company runs itself. Favourites measure attention. Modified versions measure how much people build on something. None of them measures quality, revenue or market share, and most of the benchmark scores come from the companies whose models are being scored.
Where sources conflict, I have kept the conflict visible rather than smoothing it. Two of the capability scorecards give different numbers for the same model because they were read weeks apart and the scale itself was revised in between. Counts of modified models range from 151,000 to over 300,000 depending on who is counting and what they consider a modification. Two models are reported at slightly different sizes by different outlets. The licence terms for several of the government-backed models could not be confirmed from primary sources, so they are marked unverified rather than guessed at.
Figures about the model library itself — download counts, modified versions, licence proportions, the gap between famous and used — come from Hugging Face’s summer report. Figures about the wider market — traffic share, capability scores, the developer survey, the policy timeline — come from Mozilla’s survey and the sources it cites. Prices and model-level detail come from OpenRouter and from the makers’ own documentation. The analysis and the conclusions are mine.
Primary sources
- Hugging Face — State of Open Models: Summer 2026 Observations, 14 Aug 2026: huggingface.co/blog/state-of-open-models-summer-2026
- Mozilla — The State of Open Source AI, v1.0.1, July 2026: stateofopensource.ai
- OpenRouter — The Open Weight Models that Matter: June 2026, model pages and rankings
- OpenRouter & a16z — State of AI: An Empirical 100 Trillion Token Study, arXiv:2601.10088
- Stanford HAI — AI Index Report 2026, 13 April 2026
- Epoch AI — Open models lag state-of-the-art closed models by 4 months (CC-BY)
- International AI Safety Report 2026, arXiv:2602.21012
- Mozilla / SlashData 2026 developer survey · Menlo Ventures, The State of Generative AI in the Enterprise
- European Commission — AI Act, Digital Omnibus (Regulation (EU) 2026/1744), Frontier AI Grand Challenge / EUROPA
- Model cards and launch materials: Moonshot AI, Alibaba Qwen, DeepSeek, Z.ai, NVIDIA, Meta AI Research, Google DeepMind, OpenAI, Mistral AI, Thinking Machines Lab, Black Forest Labs
- TII / ATRC — Falcon-H1 Arabic, 5 January 2026 · Counterpoint Research — Sovereign AI LLM Index H1 2026
- IndiaAI Mission, Sarvam AI, BharatGen — February 2026
- Anthropic — Fable/Mythos access statement: anthropic.com/news/fable-mythos-access
All figures current to 19 August 2026. This market moves in weeks.