Even Elon Musk surrenders to the open Chinese AI model Kimi K3. It is not for less

It’s good, it’s pretty and it’s (quite) cheap. We met him a few days ago, but Kimi K3the new open AI model from the Chinese startup Moonshot AI, is causing a sensation. So much, so much, that they have had to pause new subscriptions because they cannot handle so much demand.

Another turning point for Chinese AI. Kimi K3 is the largest open weights AI model ever published, with numbers that probably rival those of the frontier models from Anthropic and OpenAI, which do not provide information on the size of their models. Those 2.8 billion parameters make a difference and are a good part of the reason why this model represents a real leap in quality according to all the benchmarks that are being published.

“Awesome”. Elon Musk himself published a single “Impresionante” on his X/Twitter account as answer to the very complete analysis Artificial Analysis performance. Its agentic behavior surpasses that of Opus 4.8 and only Fable 5 surpasses it, but in a specific benchmark it goes even further and is the best of all the models evaluated by this firm, including those from OpenAI and Anthropic.

Artificial
Artificial

Source: Artificial Analysis.

More tests. In programming it is better than Opus 4.8 and GPT-5.5, but inferior to Fable 5 or GPT-5.6, and all the independent tests validate these results: we are facing a model that at least on paper competes directly with the best that both Anthropic and OpenAI had until now. No Chinese model had come so close until now: GLM-5.2, although notable, competed more with GPT-5.5 and Sonnet 5 than with the US frontier models.

Screenshot 2026 07 20 At 13 13 02
Screenshot 2026 07 20 At 13 13 02

Source: Artificial Analysis

Gigantic… and not so cheap. DeepSeek showed that it was possible to access really capable models at a very affordable price, and recently GLM-5.2 proposed exactly the same: it is possible to achieve 90% capacity of frontier models such as Opus 4.8, but at 20% of the cost. The curious thing is that with Kimi K3 the trend changes: it is a more affordable model than Fable 5 or GPT-5.6, but not as much as one might expect: the cost per million input/output tokens is 3/15 dollars, while in Fable 5 it costs 10/50, Opus 4.8 costs 5/25 and GPT-5.6 Sol costs 5/30.

Tokens everywhere. One of the factors that probably influences that quality/price ratio is the large number of tokens that Kimi K3 seems to use when answering. It is a model that “thinks a lot”, and that, although it undoubtedly improves the precision and capacity of the model, also causes it to generate higher bills for the user. Artificial Analysis’ own report goes further: the cost per task in its test battery is $0.95, at the level of GPT-5.6 Sol’s $1.04 and certainly cheaper than Fable 5 ($2.75), but also much more expensive than Grok 4.5 ($0.31) or GLM-5.2 ($0.47).

The pelican test. Analyst Simon Willinson was able to test the model to perform a test to evaluate the behavior of all these developments: having the model generate an SVG image of a pelican on a bicycle. In their tests the image was of very good quality, but it generated almost 17,000 tokens for the response with a task cost of 25 cents. It is not that this test is too conclusive, but it does reveal that for a simple task, the result, although outstanding, is not especially efficient in token consumption.

Cybersecurity, the unknown. Unlike the latest models from Anthropic or OpenAI, Moonshot AI does not seem interested at the moment in its use in the field of cybersecurity. There is no mention of those potential capabilities in the notes of launch, but that doesn’t mean it doesn’t deliver. Vercel’s CTO, Malte Ubl, explained Although it is not the most advanced of AI models in this area, after running several tests it seemed like a model that can be very useful when finding and correcting vulnerabilities.

Demand, through the roof. The expectation generated by this model has been such that the company has announced that pause new subscriptions. This will allow them to be able to deal with all requests to use it without harming the experience for both old and new users. A striking decision that seems to make a reality clear: they cannot cope.

In Xataka | The gigantic Qwen 3.8 is another worrying sign for the US: its AI advantage is evaporating

Leave your vote

Leave a Comment

GIPHY App Key not set. Please check settings

Log In

Forgot password?

Forgot password?

Enter your account data and we will send you a link to reset your password.

Your password reset link appears to be invalid or expired.

Log in

Privacy Policy

Add to Collection

No Collections

Here you'll find all collections you've created before.