← newsletter

Agentics: Thoughts on Open Models and Kimi K3

Amol Kapoor · July 21, 2026

I.

Things in AI change so rapidly that you’d think nothing would surprise us at this point, but this past week the AI world was rocked by the release of Kimi K3, a new open weight model coming from the Moonshot lab in China. This is the world’s biggest open weight model, clocking in at 2.8T parameters. It is also, by every account, very very good.

I’ve been running Kimi K3 alongside Claude on my normal coding work, and for all practical purposes I can’t tell them apart. Same tasks, same quality of output, and near identical token counts to get there. I expected an open model to be sloppier or to grind through more tokens on the way to the same answer, and neither turned out to be true. — The Kimi K3 Moment

On our private long-horizon knowledge work evaluation, Kimi K3 reaches an overall Elo of 1547, +732 points from Kimi K2.6 and behind only Claude Fable 5…Cost per task ($0.94) is similar to GPT-5.6 Sol ($1.04), ~1/2 the price of Opus 4.8 ($1.80) and higher than open weights peers — Artificial Analysis

Artificial Analysis Intelligence Index showing Kimi K3 among the leading models.

On the first try, Kimi K3 just found the source of a bug that Fable 5 hasn’t been able to pinpoint in multiple attempts. It’s just one anecdote, and I haven’t used K3 much yet, but so far it’s looking extremely promising. — HN

Axios maybe put it the most bluntly with this headline:

Axios headline: China just erased America’s AI lead.

Perhaps the most impressive demonstration to me was this incredibly detailed replica of the entire MacOS operating system. Supposedly some guy had a swarm of Kimi instances making this over the course of a few days. It’s kinda incredible, you can go in and edit files, use a fully functioning terminal, listen to music, check emails. You can go to the ‘App Store’, ‘download’ the ‘Chess app’, and then play chess (no quotes around that one, you can actually do that!)

Kimi-built MacOS replica running a chess app.
I have a special opening that is a guaranteed winner.

One popular way to conceptualize these models is that they are released as part of coherent generations, which are defined by the time of release and the capability of the model. If you plot models along these two axes, you get some clear clusters along the frontier.

Plotting model release date against model capability as measured by Artificial Analysis. Grouped into generations entirely based on my vibes.
Generation One model capability cluster led by GPT-4.
Generation Two model capability cluster including Claude 3.5 Sonnet, GPT-4o, and Gemini 1.5 Pro.
Generation Three reasoning-era model capability cluster.
This era is interesting because each individual player managed to hold the top spot for a few months. It is also the only era where Google really had a dominant position. RIP Gemini 2.5, you were my favorite for a while.
Generation Four agentic-era model capability cluster.
Generation Five model capability cluster with open models moving to the frontier.

There are a few trends that are worth pointing out.

First, even though progress in AI is already progressing at an incredible clip, somehow things are still accelerating. New model generations seem to be releasing at a faster and faster pace, with the time between releases compressing.

Second, the rate of acceleration is not consistent across organizations. The frontier continues to be pushed forward, but the labs that are at that frontier continue to shift. This is most obvious when you look at the trajectories of Copilot — once the only competitor to ChatGPT, now widely derided — or the Meta series of open source models, which were surpassed by the Chinese models a year and a half ago and never really recovered.

Third, the rate of acceleration is not consistent across countries. In relative terms, the US continues to be in the lead when it comes to raw model capabilities. The French (Mistral) were at one point quite competitive but have since fallen off. And, most relevant to this piece, the Chinese labs have continued to forge ahead, rapidly closing the gap against the top US labs.

Despite the above, up until last week, it was somewhat unfathomable to believe that anyone would have a better model than the top US labs. Sure, some folks were saying that certain policy decisions may result in the Chinese labs taking the lead, but this was strictly in the realm of the hypothetical.

Kimi K3 blows all of this out of the water.

It was only 6 weeks ago that the US federal government was banning ‘Mythos-class’ models for being too dangerous. Now, we have a ‘Mythos-class’ model that is not just outside of US control, but is in fact open source and available to anyone in the entire world to download and use. To add insult to injury, Kimi K3 is significantly cheaper to run. Hosted Kimi models cost a fraction of the equivalents from Anthropic and OpenAI ($3 vs $10 and $5 per million). Meanwhile, you can get a lot of juice with the ~$40 kimi subscription, way more than anything Anthropic or OpenAI are offering in the same tier. The max kimi subscription of $100 per month is about equivalent to a $200 Claude Max sub. In a world where people are becoming ever more cost conscious around their token spend, the Chinese labs are less ‘also ran’ and more ‘default option.’ And the Chinese labs are fighting with a handicap! They are running on significantly worse hardware due to ongoing export restrictions of high end silicon chips!

The markets are reacting with surprise. Kimi’s release triggered sell-offs of Google and SpaceX, in addition to the wider market of AI stack stocks. I guess I don’t know why they’re surprised, in some sense this is a long time coming. Do you all remember Deepseek? I wrote about it a year and a half ago when it came out, and my analysis then wasn’t so different from my analysis now:

To be clear, the fact that Deepseek exists isn’t really that significant. Everyone always knew there was going to be a big Chinese model. The country is not afraid to wall off it’s populace from the West; it was never going to allow LLMs that are happy to tell people about Tienanmen Square and Winnie the Poo and Taiwan. So when the first Chinese LLMs came on the market, everyone was like, “Yea, whatever”. The first batch, the Qwen models from Alibaba, were, like, fine.

Even though I expected a set of LLMs to arise thanks to protectionism and state-interest, I (and everyone else) assumed those models were just going to be worse than the US ones. This is how technical development has always been, after all. Google Search is good, Baidu is eh, and if you can use Google over Baidu you do. Amazon is good, Alibaba is eh, and if you can use Amazon over Alibaba you do. Facebook is good, Tiktok is eh, and if you can…no wait that doesn’t work does it?

Anyway, the larger point is that no one really thought the Chinese models were a threat. Sure, people would talk about how the Chinese government was a threat, but it was always in a hypothetical way, mostly used to justify infinite capitalist investment without any corresponding concerns about AI safety or alignment. I don’t think most people actually thought that a Chinese company would come out and deploy a model that is simply better than what we have in the states.

Clearly we didn’t actually learn anything from TikTok.

And we still aren’t learning. Even now, some people are clearly just in denial. I’ve heard several folks say some variant of “it’s easy to catch up when you’re just stealing / distilling from the bigger labs,” (more on distillation below) as if this somehow negates the fact that one of the best models on the market is not US-made.1 The conventional wisdom is no longer relevant. There are no secret models hiding in the backrooms of OpenAI / Anthropic. The Chinese models are no longer 6 months behind, they are at par.

Kimi’s release has had me rethinking some of my previous position on LLM economics. A month before Deepseek was released, I argued that the LLM market is a bit like the search market in that there are clear winner-take-all dynamics.

LLMs are pretty easy to make, lots of people know how to do it — you learn how in any CS program worth a damn. But there are massive economies of scale (GPUs, data access) that make it hard for newcomers to compete, and using an LLM is effectively free so consumers have no stickiness and will always go for the best option. You may eventually see one or two niche LLM providers, like our LexusNexus above. But for the average person these don’t matter at all; the big money is in becoming the LLM layer of the Internet.

The economics of LLMs means that it is critical for these players to have the best models. There’s no room for second place.

A year and a half later, it’s clear that I was wrong about a few things. First, I was wrong that the LLMs are effectively free. They aren’t, as we saw with the rise and fall of tokenmaxxing. The cost per unit intelligence has dropped precipitously, but the overall cost per frontier token has skyrocketed. Because I was wrong about the pricing, I was also wrong about the quality / cost curve. There is room in the market for multiple models at different price points, because consumers can in theory choose different providers for different tasks. The rise of model routing as a service and the increasing classification of tasks and roles into different AI usage categories — sales gets a cheap open source model, engineering gets fable — is downstream of that demand.

But I was really right about the most important bit, which is that consumers have no stickiness at all.

Every three weeks we see a mass exodus from Anthropic to OpenAI to Anthropic and back to OpenAI. It is just way too easy to switch providers. “Model wrapper”, once seen as derogatory, is now touted as a core feature. Businesses and consumers are recognizing both the churn and the need to stay on top, and are explicitly investing in tools that let them adapt. This is also a huge part of why the background agent infrastructure that we build and sell at Nori has zero lock-in at the compute, model, and harness levels — there is no reason to be locked into a single ecosystem when you don’t have to be (if you’re looking for background agents / cloud agents for your team, shoot me a message!)

The moment the Chinese models hit the actual frontier, there will be a mass exodus to using those models. The economic incentives are way too strong, and arguably that is already happening for people who are in the know.

It’s undeniably true that the Chinese labs did train on reasoning traces and outputs from Claude et. al. But also, Kimi outright beats Claude et. al. on several benchmarks, which is unlikely for a raw student-teacher training paradigm

II.

This is not how any of this was supposed to go. The whole point of all this closed model stuff is the ability to dig a trench around the model and put a toll booth in the middle, and the whole point of all the debt is to basically have a call option on that toll booth. We all knew that the big labs would be fighting tooth and nail with each other over which of their closed models would end up winning the day. Doing this open source stuff feels against the rules.

Partially, that’s because it is against the rules. Many of the open source models are trained on their more capable closed source counterparts.

Unfortunately, usage of the models is a source of valuable training data for competitors. Every improvement to any model can quickly be copied by other labs even without seeing any of the internals because you can just sample the new model a bajillion times and use the outputs as training data. This technique is called ‘distillation’ because you are ‘distilling’ the essence of some larger ‘teacher’ model into a smaller ‘student’.

OpenAI and Anthropic can try to ban people from training on the outputs of their models, they can kick and whine and sue to try and enforce it, but at the end of the day the labs need people to use their models. That’s the foundation of all of the token economics! If no one actually uses your models, what’s the point of all that training?

So the token traces have to get out in the wild, and any kind of adversarial cat and mouse is going to end up wasting a lot of resources for very little gain. The open source models will basically always be able to catch up, modulo some amount of compute.

Kimi K3 responding that it is Claude when asked what model it is.
Kimi K3 identifies itself as Claude, likely a reflection of the sheer amount of Claude-output in the training data.

This isn’t a new problem, people have been talking about it for some time. Dwarkesh even asked folks to tackle this problem in his essay competition last month:

So when does the profit start? Maybe at some point scaling will plateau, but if progress at the frontier has slowed down, then the combination of distillation and low switching costs (cloud margins result from high switching costs) makes it really easy for open source to catch up to the labs, eating into their margins. So how do the labs actually start making money?

(emphasis mine)

My answer was that the labs in totality will not be able to meet the demand for their models and will be capped by compute, which in turn will prevent new-comers from actually competing for best in class model training and will let some of the labs rent seek. The winning answer argued that the models will become commodity and the model providers should try and own downstream services instead. Note that both of these answers assume that distillation is so inevitable, it’s hardly worth discussing.

This has already had some amount of impact on OpenAI and Anthropic’s pricing models. Anthropic, for example, continually extends the amount of time that Fable will remain on its subscription plans. This is not entirely attributable to the open source models, but I am certain that these models will add additional pressure to simply keep Fable on the subscriptions forever — even if doing so results in a loss for Anthropic, since the subscriptions are heavily subsidized.

III.

I don’t mean to take away from Moonshot’s accomplishments. Kimi K3 is an incredible technical achievement on its own terms. The hardware limitations have forced the Kimi researchers to pull some pretty neat tricks, all of which are now open source and can be adopted across the ecosystem. Kimi also forces both OpenAI and Anthropic to keep their prices down, which I (as a startup founder) am a huge fan of.

The model weights themselves have not yet been released — they are slated to drop on the 27th. When they do, I will be very eager to pull Kimi into our background agent environments and take it for a spin.

The US is no longer the only nation capable of frontier intelligence, no longer in a category of its own. The world is a smaller place now.

Adapted from 12 Grams of Carbon.