DeepSeek CEO explains how he has transformed the chip shortage into his strategic alibi

One of the many battles that China and the United States are fighting is technological, specifically one related to artificial intelligence. US business spending on AI It’s astronomicalbut in China the ambition is no less and the Government wants it to be one of the pillars for the country to be the first world power in the short term. One of the conversations surrounding Chinese AI compared to American AI has to do with efficiency. Very capable models, higher in some cases to the Americans, but with a much lower price and, above all, trained in record time. There are controversies about it, such as theft accusations by giants like Anthropic and OpenAI, but the reality is that if you don’t want to use the AI ​​of an American company, they are there alternatives like DeepSeekKimi or those from Zhipu AI. But, speaking of Deepseek, the company’s CEO is clear that there is a Huge difference between Chinese AI Big Tech and American counterparts: computing capacity. He gives a fact to illustrate it: four Huawei chips are equivalent to one from Nvidia, and that opens the door for the US to have even more arguments to continue putting pressure on the Chinese technology industry. The bottleneck is power, nothing more Liang Wenfeng He is the CEO of DeepSeek and, like so many in his position in recent weeks, he has had the presentation before investors. These types of events are very interesting because they allow us to learn more about both the companies and their plans, and in this case it has been revealing to listen to the head of one of the most relevant AI companies today. They are not official statements, since they are the leak of the almost four-hour meeting published by Tencent Tech, but among its many phrases, we can extract very interesting ideas. The first is what we mentioned: for the boss, the main gap is money and resources. “There is no gap in personnel, since it is basically the same group of people. Talent is not the bottleneck: resources are the biggest bottleneck,” says Liang. The CEO continued to point out evidence, stating that “more cards are always better” and that, at a reasonable price, they buy as many as they can because “converting money into Nvidia cards is definitely better than leaving that money in the bank.” And the Nvidia cards stand out for a fact that is also very interesting: how many Huawei cards are equivalent to one Nvidia card. “Our price offers a reasonable profit: we buy a batch of equipment and they pay for themselves in about ten months” – Liang Wenfeng Assures that the relationship is four to one, or what is the same: if they wanted to train a model with 800,000 million active parameters like the main American AI models, they would need 50,000 Nvidia GB300 cards or 200,000 Huawei 950 cards. That It doesn’t seem to leave Huawei in a very good place.which are those that the Chinese industry and the Government itself is trying to promotebut Liang affirms that Nvidia is digging its own grave in the Chinese market, that the CUDA ecosystem is eroding quickly and that Huawei supernode 950 It can completely replace Nvidia’s GB200 and GB300 in both performance and price. Although the 4:1 ratio between the chips seems alarming for the domestic company, Liang did not specify the workload or whether it measures raw performance, training performance or inference. The US looks closely One detail that Liang points out is that this computing power gap will end up closing as Huawei launches solutions and that, although he does not dare to guess when it will happen, it is something that will end up arriving. But of course, when the competition between the US and China is, in part, in that power of AI, the United States is not in the least interested in closing the gap. In recent months we have seen how pressure has continued to prevent Chinese companies from accessing Western hardware that is key to the development of the semiconductor industry. ASML Extreme Lithography Machines are an exampleand more recently the White House accused Chinese startup Moonshot AI from acquiring servers equipped with Nvidia chips that they should not be able to access. The case of Moonshot AI has been quite popular because a few days ago we told you about that amazing Kimi K3 model that, according to the Americans, would have been trained in Thailand with Nvidia GB300 chips to avoid United States export regulations. And they go further, arguing that Kimi K3 had ‘drunk‘the model Anthropic Fable. Deepseek Strategy Beyond politics, and returning to Liang, although right now it seems that China is not on par in terms of computing power, DeepSeek has a long-term strategy that has less to do with having the most powerful model currently. As we read in SCMPLiang commented that they are setting prices “to make only a reasonable profit, not to maximize revenue.” And that is the key that the company apparently pursues, since they are keeping prices low for users with the aim of their service being increasingly used in the hope of increasing the possibility of achieving artificial general intelligence. So while other companies prioritize immediate profits and capturing market share, Liang is focusing on “increasing the probability of achieving AGI“because it will be that AI that will have great commercial value. And that is what they do seek to exploit. In Xataka | While most oppose AI data centers, there is one group enthusiastic about them: merchandise thieves

China has a plan to win the AI ​​war against the US. And DeepSeek is its champion

Liang Wenfeng is the most elusive person in the Chinese manufacturing industry. artificial intelligence (AI). The founder of DeepSeekwho also runs the hedge fund High-Flyer, recently held a four-hour video call with potential investors from Hangzhou, China, an unusual format in every sense: only two representatives per institution, and for most of the attendees it was the first time they had seen the founder. “We are a very normal group of people,” Wenfeng told them. And this apparent simplicity hides the company that, in just over a year, has rewritten the economic rules of generative AI. However, what began as a start-up reluctant to any outside investment has come full circle. DeepSeek recently closed a financing round of around $7 billion, with a valuation of $52 billion. The most revealing data, however, is provided by SCMP: This company is already negotiating a second round that is expected to raise that figure to $71 billion just weeks after closing the first. The power map of Chinese AI Reuters has confirmed that DeepSeek develops its own ASIC chip aimed at inference and not training. With it, it aims to reduce its dependence on Nvidia and Huawei, its two current suppliers. If the project succeeds, it would mark a huge strategic shift for a company that until now has built its entire reputation on software efficiency rather than hardware control. Be that as it may, this move would add additional pressure to Huawei, which competes for the same space within China. Huawei now integrates DeepSeek into its cloud services in sub-Saharan Africa In this scenario we are interested in focusing on the real magnitude of this phenomenon. According to FortuneChinese open source models (led by Qwen, MiniMax and DeepSeek) already represent a third of global use of large language models, compared to an almost non-existent presence at the end of 2024. Emerging companies in Silicon Valley and Southeast Asia adopt them due to their openness, transparency and a much lower operating cost than American alternatives. Huawei, in fact, already integrates DeepSeek into its cloud services in sub-Saharan Africa. As expected after such an escalation in valuation, several sources suggest that DeepSeek could present its IPO this year, following in the wake of other low-cost Chinese startups such as Zhipu AI and MiniMax, which are already listed on public markets. A successful placement would give DeepSeek the institutional capital necessary to scale its computing infrastructure and sustain its long-term price war against large American players. DeepSeek is no longer just the startup that sparked Nvidia’s biggest stock scare in a single day: it is the battering ram with which China tries impose its open and cheap AI model as a global standard. Image | Generated by Xataka with ChatGPT More information | Bloomberg In Xataka | While most oppose AI data centers, there is one group enthusiastic about them: merchandise thieves

DeepSeek no longer wants to compete only with models. Its new front aims directly at NVIDIA’s business, according to Reuters

In just over a year, DeepSeek has stopped sounding like a rarity in the Chinese industry to become one of those names that already appear every time we talk about the global race in artificial intelligence. First we look at it for its models, for its efficiency and for the shock it caused beyond China. Now the question begins to move to another terrain: what happens when a company that competes in software understands that the next advantage may be in the chips that make it possible to execute that AI on a large scale. The jump to hardware. The information that opens this new front comes from Reuters. The agency assureda, citing three people familiar with the matter, that DeepSeek is developing its own artificial intelligence chip, aimed at inference tasks and not training new models. We will see the technical nuance immediately, because it changes the reading of the movement quite a bit. For now, caution is mandatory: DeepSeek has not publicly confirmed the project would be in an early phase and the company did not respond to the agency’s request for comment. The key is in inference. The easiest way to understand this is to think about what happens after training. Once the model is built, every question we ask and every answer we receive requires putting it to work again. It is not an isolated operation, but a routine that is repeated millions of times if the product works. That is why a chip designed for that phase does not aim so much at technical prestige as at something more earthly: making using AI cheaper, faster and less dependent on third parties. The move is best understood if we look at what DeepSeek has depended on so far. The company has used chips from NVIDIA and Huawei to train and run its models, including the base that held R1, trained on NVIDIA H800a chip designed for the Chinese market whose export to China was banned by Washington at the end of 2023. Since then, DeepSeek has increasingly relied on Huawei: In April it launched its V4 model adapted to Ascend and Huawei said its processors were used in part of V4-Flash training. DeepSeek is no longer a footnote: Until not so long ago, the global debate on AI seemed to revolve almost entirely around American companies such as OpenAI, Google, Microsoft, Meta or Anthropic. DeepSeek changed part of that conversation by demonstrating that China could also produce models capable of circulating outside its domestic market and force the industry to look towards Hangzhou. Recall that the company was widely celebrated in China as a national AI champion. The trend is already seen in a good part of the sector. Google has been developing its TPUs for years, Amazon has Inferentia for inference payloads, Microsoft has Maia and Meta works at MTIA. Reuters also cites two recent movements especially close to the case: OpenAI announced its Jalapeño chip with Broadcom in Junealso oriented to inference, and Anthropic was considering designing its own chips. The pattern is quite clear: large AI companies want to rely less on third-party providers and better control the cost, performance and availability of the computing that powers their services. The big obstacle is manufacturing it. Designing a competitive chip is not the same as wanting to have it. Developing an AI accelerator typically requires years, a lot of capital, and a network of design, foundry, and memory partners. For a Chinese company, furthermore, the problem does not end at the technical level: US export controls limit access to the most advanced foreign factories and also to high-bandwidth memory, a key component for this type of chips. Times change. NVIDIA arrived at the AI ​​boom with an advantage built over decades: in 1999 it launched the GeForce 256, presented by the company itself as the industry’s first GPU, and in 2006 launched CUDAthe architecture that helped take the parallel processing of its chips beyond graphics. When the models started requiring massive amounts of compute, I already had the hardware and ecosystem in place. For years, for much of the industry, competing in AI meant going through its chips. What the DeepSeek case suggests, with all caution, is that this dependency is beginning to have cracks. Images | Xataka with Nano Banana In Xataka | Samsung earns 19 times more than a year ago. Investors have reacted by sinking the stock 7%

DeepSeek is good, pretty and very cheap. And above all, the weapon to create a Chinese hardware industry independent of Nvidia

The arrival of DeepSeek-V4-Pro It hasn’t caused that much of a stir. like the one caused by DeepSeek R1 a year and a half ago, but we may be facing an even more important model. If that version revealed to the world that China was advancing spectacularly in this race, this other one is beginning to allow us to glimpse something else more interesting. What most people see is a very decent model and above all “low priced”. Which hide the company It’s another more important thing: achieve independence from Nvidia and US hardware. what has happened. Last Friday, those responsible for DeepSeek announced something surprising: their promotional offer with a 75% price cut to use their DeepSeek-V4-Pro model will be maintained permanently. That makes this model offer very decent features (but not exceptional) for a really low price: 1M entry tokens 1M tokens output DeepSeek-V4-Pro 0.435 0.87 GPT-5.5 5 30 Opus 4.7 5 25 Gemini 3.5 Flash 1.5 9 Good, pretty and very cheap. It is true that the performance of DeepSeek-V4-Pro is inferior to that of rival models from OpenAI, Anthropic or Google. Artificial Analysis tests indicate that the DeepSeek model is at a very good level, but it is also much cheaper than its competitors. This is especially relevant for agentic tasks that consume many tokens and that with this model become accessible and very affordable. According to Artificial Analysis, DeepSeek is close to the performance of the best models in the industry, and although it is slower in its responses, it is also much cheaper than the frontier models from OpenAI, Anthropic or Google. A different strategy. How is this company going to make money? It does not have subscription plans like its local competition (GLM, Kimi) or the western one (ChatGPT Plus, Claude Pro). It also does not have voice or image models. It does not have an AI agent for programming that competes with Claude Code. It publishes the open weights of its models and shares its technical innovations with the industry (and with its competitors). For those who closely follow the company and these decisions, the strategy is clear. DeepSeek’s goal is not to win the AI ​​model race. Their goal is to build a Chinese AI hardware industry that doesn’t depend on Nvidia or TSMC… and get paid their share in that process. Hardware independence. China has a structural problem in this AI race: sanctions and vetoes imposed by the US make you unable to access the most advanced chips nor to ASML UVE photolithography. And since China cannot currently compete in terms of computing power, what its companies are doing is ensuring that their AI models need less computing power to achieve similar results. Efficient architectures. The Mixture of Experts (MoE) and Multi-head Latent Attention (MLA) architectures are two key weapons in this strategy. The first already existed but was adapted by DeepSeek for their model: with it only part of the total parameters of the model are activated to answer the query without losing precision. What MLA does is compress the attention information (the so-called KV Cache) with which the model maintains the context of a conversation, reducing it by 90%. Both techniques allow us to reduce the need to use high-speed HBM memories, something that is also striking in order to reveal DeepSeek’s probable strategy. The importance of KV Cache. As the GDP analyst explains in Xthat use of MLA allows that for one million tokens, DeepSeek-V4-Pro only needs 5.48 GB of HBM memory. Competitors like Zhipo AI, which develops GLM 5, need 60 GB for the same, while Alibaba’s Qwen 3 needs 89 GB. This advantage allows DeepSeek to offer much lower prices to obtain performances similar to those of its competition, but it also means that DeepSeek models can run on Chinese memory chips that cannot compete in speed with HBM modules. Goodbye HBM, hello NAND and SSD. These innovations open the door to the use of NAND memories and even SSD drives to process this data, and there YMTC enters the scenea Chinese Flash memory manufacturer that is slowly becoming a global giant. Also CXMTwhich manufactures DRAM memoriesbecomes an alternative here and the reason is equally interesting: DeepSeek introduced a memory search module in LLMs called Engram which is also intended to avoid excessive dependence on HBM memories. How to bypass the CUDA monopoly. Nvidia continues to have a fundamental element in CUDA to maintain its market dominance, but here DeepSeek too has proposed an alternative. Is called Tile Kernels and these are software cores created with TileLang (a variant of Python for this field) that allow governing advanced AI chips (GPUs). Huawei as an invisible ally. Those responsible for Huawei recently indicated that its new Ascend AI supernodes fully support DeepSeek v4 models. Precisely this provides another fundamental advantage to the company, which thus avoids (at least in part) total dependence on the use of Nvidia chips and prepares to further strengthen Huawei’s relevance in a market in which until recently Jensen Huang’s company was queen and mistress. Open models to attract the hardware industry. US companies continue to maintain their closed and proprietary models, but DeepSeek is one of the many Chinese startups that publish them with open weights. With this, what she and the others intend to do is not only attract AI developers and users, but also create a hardware ecosystem that adopts these architectures. DeepSeek invites its rivals to use techniques such as MoE or MLA precisely so that all these advances become a de facto standard and hardware manufacturers also adopt them and integrate them in an optimized way into their designs. A round of 10,000 million to advance. The company is also preparing a financing round in which they intend to raise 10,000 million dollars and with which they would achieve a valuation of between 45,000 and 50,000 million dollars. Still far from the mammoth valuations of OpenAI or Anthropic (already close to a billion dollars) but certainly … Read more

DeepSeek V4 has given China the boost it needs against the US. Four chip makers are the big winners

DeepSeek V4 It is the catalyst China needed. This model of artificial intelligence (AI) developed by the quantitative hedge fund specialized in trading algorithmic High-Flyer has been designed natively to live with Chinese chips. This is exactly the strategy that the Chinese government supports in response to the pressure that the US is putting on China. The Administration led by Donald Trump prevents the most powerful GPUs from Nvidia, AMD or Cerebras from reaching this Asian country. And Beijing has decided to do without them. The challenge facing the Chinese government is that it is much easier to set this goal than to put it into practice. This is the scenario in which DeepSeek V4 has emerged as the asset that China needs. And its arrival has led, for the first time, to several Chinese AI chip designers achieving something that until now had only been within the reach of Nvidia: guaranteeing full compatibility with the latest High-Flyer AI model from day 0. A great opportunity for Huawei, Cambricon, Moore Threads and Hygon DeepSeek V4 has marked a turning point. Its adoption in China is likely to be very notable, which has caused AI chip designers to compete among themselves to ensure full compatibility with this model from the moment it arrives. None of them wants to miss the opportunity to grow in the largest market on the planet if we stick to the most relevant indicators, such as purchasing power parity or the volume of population with the capacity to consume. Huawei is one of the companies that benefited most from the arrival of DeepSeek V4 Huawei will surely be one of the companies that will benefit the most from the arrival of DeepSeek V4. And its entire portfolio of GPUs for AI is compatible with this model. Nevertheless, your Ascend 950PR chip has been established as the main inference solution. A note before moving forward: inference is broadly speaking the computational process carried out by language models with the purpose of generating responses that correspond to the requests they receive. China’s three largest internet groups (Alibaba, ByteDance and Tencent) have placed orders for several hundred thousand Ascend 950PR processors following the launch of DeepSeek V4, according to Reuters. However, Huawei is not the only Chinese company that has won the lottery with the arrival of this AI model. Cambricon Technologiesthe Chen brothers’ company, has already completed the adaptation to the framework open source vLLM inference framework and has published the code on GitHub. Besides, Moore Threads has worked closely with the Beijing Artificial Intelligence Academy to run DeepSeek V4 on its MTT S5000 card using the FlagOS software stack. And Hygon has carried out a deep optimization of this model in its DCU platform with the purpose of consolidating its hardware as an attractive option for industrial use. The competitiveness of DeepSeek V4 outside of China is unclear because is less capable than its more advanced American competitors, but its future within the borders of its home country appears to be guaranteed. Image | Huawei More information | SCMP In Xataka | The US’s problem in the AI ​​and humanoid race is not China: it is all of Asia and it is greatly disadvantaged

DeepSeek wants to raise its first round of financing and copies the last thing that remained to be copied from the US: the economic model

Chinese AI startups appear to have surrendered to Silicon Valley capitalism. Both DeepSeek such as Moonshot AI (Kimi) have begun to raise investment rounds or are preparing to do so. It is a turning point in a race that is now becoming especially interesting and that also raises a clear question: will these companies continue betting on open models? The valuation is multiplied by two. DeepSeek had always avoided making that decision and it seemed almost a personal project of its founder, billionaire Liang Wenfeng. However, the company is now in talks to raise its first round of external investment, they assure in Financial Times. According to company data, Wenfeng has 89.5% of the stake in the company. There is talk of a round that would increase DeepSeek’s valuation from the current $20 billion to around $45 billion. Who is the “Big Fund”. Behind this investment round is above all the China Integrated Circuit Industry Investment Fund, also known as the “Big Fund”. This consortium, the most important of its segment in the field of semiconductors, is supported by the Chinese state, and has a “cash” of 47 billion dollars contributed by the Chinese Ministry of Economy, the local government and several state banks thanks to a third round that was carried out in 2024. At the moment the “Big Fund” has not invested in other Chinese AI startups, but it has in companies like SMIC or Yangtze. The war for talent. The reason behind this decision is not only the need for capital to have access to more computing capacity. According to sources close to the operation, Liang Wenfeng has been forced to open that option to stop talent theft and thus be able to keep their best researchers on the payroll. In a market as competitive as this one, DeepSeek needs to offer shares to its employees to compete with the aggressive recruitment of talent by its local and Western rivals. A promising pairing. The relevance of this investment goes beyond the AI ​​model as such. DeepSeek has been significantly optimized for be able to run on Huawei hardwareallowing China to have a platform that works without the need for Nvidia chips. This symbiosis between this efficient AI model and the Chinese hardware giant is quite a bet by the Chinese government to try to win this race despite Washington’s blockades. The forced bet on “national” chips. Seeking that support in Huawei chips is not only a technical choice, but a political necessity for survive NVIDIA GPU crash. The problem is that Chinese hardware is still struggling to close the raw performance gap against architectures like Blackwell’s. If DeepSeek’s software hits a ceiling and chips created in China do not evolve at the necessary pace, the laboratory could find itself trapped: it would not matter to be very efficient when they cannot compete in raw power. Moonshot signs up for the rounds. DeepSeek is not alone in this race to achieve huge valuations. Moonshot AI just got up 2 billion dollars from investors such as Meituan, raising its value above 20 billion. Meanwhile, other rivals such as MiniMax and Zhipu AI (GLM) already surpass the 30,000 million valuation in their stock market debuts. This trend is therefore following what was already experienced (and continues to be experienced) in the US with AI startups, and the capital bubble that exists in the North American country now seems to have its eastern version in China. Moonshot AI and exceeds $200 million in annual recurring revenue (ARR). The paradox of copying the economic model. It’s ironic that DeepSeek, which became famous for challenging the “brute force” of American spending, ends up adopting its same funding structure. The company has shown that efficiency could offer an alternative to those almost unlimited resources of venture capital accessed by OpenAI or Anthropic. However, market reality dictates that a very solid capital structure is still what is needed to survive in the long term. Either you have it, or you can’t continue training models, reserving computing capacity and, of course, retaining talent. Open models? Until now DeepSeek had been one of the heroes of open weight AI models. Thanks to this, platforms like Hugging Face allow you to download it and allow everyone to take advantage of its achievements in terms of efficiency. The entry of venture capital and state funds could change the rules of the game: investors do not usually inject billions of dollars so that the product ends up being “given away” even for its competitors. The company will probably face the dilemma of closing its next models to protect its valuation and generate exclusive income, or keep its philosophy open at the risk that its investors no longer trust that strategy. In Xataka | If at some point NVIDIA has to choose between giving its best chips to the US or China, its choice is very clear.

DeepSeek V4 is here. It’s good news for efficiency and bad news for the myth

DeepSeek has published its V4 model under MIT license, with notable improvements in code and architecture designed for Chinese chips. It has also admitted, in its own technical report, that it is three to six months behind leading Western models. For a laboratory that A little over a year ago the global narrative of AI changedthat is much more than a nuance. Why is it important. DeepSeek became a symbol in January 2025. Its moment shook the markets, questioned the logic of the American technology stock market and convinced half the world that China could compete head-to-head on the frontier of AI, at a fraction of the cost. It’s not that V4 destroys that story, but it does complicate it a bit. China’s most important laboratory in AI arrives with a model that its own engineers describe as a step, not a leap. The context. V4 has taken longer than expected to arrive. According to sector sources collected for 36KrDeepSeek suffered a serious training failure in mid-2025 while trying to migrate its infrastructure from NVIDIA to Huawei’s Ascend chips. Internal opinions on technical direction were not aligned, and the founder, Liang Wenfengimposed conditions that were difficult to execute. The result: months of delay and a model that, furthermore, is still not multimodal, postponed due to lack of computing capacity and cash. Between the lines. The most interesting thing about V4 is in its architecture. The model introduces TileLang, a domain-specific language that allows low-level code to be decoupled from CUDA (the NVIDIA standard) and compile it for different chips. It also incorporates MegaMoE, a kernel designed to reduce latency in expert parallelism that already runs on Ascend hardware. But V4 training has continued using NVIDIA GPUs. Independence is, for the moment, more of an aspiration than an accomplished fact. turning point. While DeepSeek looked inward, the Chinese market has been reorganizing itself without it: Doubaofrom ByteDance, has become the most downloaded chatbot in China. MiniMax and Z.ai They have gone public. Alibaba has achieved great adoption thanks to vertical applications. DeepSeek never wanted to build a consumer product, and the market hasn’t waited for it. The internal bill has also arrived: the laboratory has lost key talent to Tencent, ByteDance and Xiaomi in practically all areas. Liang Wenfeng refused to give up 20% to an unidentified large investor. And now, for the first time, DeepSeek opens an external funding round. Main loser? The narrative of open source Chinese as a real alternative to the Western closed model has taken a hit. A Qwen employee has told 36Kr that “the golden age of nonprofit AI development is over.” The big question. It’s whether DeepSeek can regain lost ground. That depends largely on Huawei, whose Ascend 950 promises to scale well with V4, but 750,000 units are equivalent, adjusted for quality, to a week of American production. The gap is not closed with ingenious architectures. It is closed with silicon. In Xataka | Companies around the world face an irresolvable dilemma: either they are with China or with the US, with both it is no longer possible Featured image | Solen Feyissa

DeepSeek has just released a model that competes with Opus 4.6. It costs seven times less and runs on Chinese chips

They have passed 484 days since that “DeepSeek moment“, but the wait It seems to have been worth it, because we have the new DeepSeek V4 with us. We are facing an absolutely gigantic open weights model that once again promises to crack the foundations of the proprietary foundational models of Anthropic, OpenAI or Google. This is moving, gentlemen. Gigantic and open. DeepSeek v4 is an Open Source model and comes in two versions. The first is the Pro, with 1.6 trillion parameters (1.6T), of which it has 49,000 million active. The second is Flash, with 248,000 million parameters (248B, huge for a “Flash” model) of which 13,000 are active. More efficient than ever. Both versions they make use of a Mixture-of-Experts (MoE) architecture, which means that only a fraction of the parameters are activated in each inference. This allows the computational cost to be reduced significantly. Both versions support a context window of one million tokens—to include novels and novels at once as input—when in v3 it was 128,000 tokens. Furthermore, this model is much more efficient than its predecessor in computing per token: it requires only 27% of the operations per token and 10% of the KV cache compared to DeepSeek v3.2. Benchmarks promise. DeepSeek’s internal testing reveals that v4 Pro-Max (the best model with the highest reasoning ability) outperforms or is on par with Claude Opus 4.6 Max, GPT-5.4 xHigh, Gemini 3.1 Pro High, Kimi K2.6 and GLM 5.1. The results, however, are not independently verified, which means we should take them with caution. The numbers are still striking: in LiveCodeBench, a programming test, DeepSeek v4-Pro-Max achieves a 93.5% score compared to 88.8 for Opus 4.6 and 91.7% for Gemini 3.1 Pro. In other tests there is more variability, but at least on paper DeepSeek v4 Pro seems as good as Opus 4.7, which until now was the absolute benchmark. Much cheaper. But as happened with its previous version, the difference in price with those models from US companies is astonishing. As point the analyst Simon Willinson, the official prices of DeepSeek v4 Pro are 1.74 dollars per million input tokens and 3.48 dollars per million output tokens, up to almost seven times less than those of Opus 4.7 and up to almost 9 times less than those of the new GPT-5.5. With DeepSeek v4 Flash the cost is 0.14/0.28 dollars per million input/output tokens, when GPT-5.4 Mini costs up to 16 times more. The conclusion is obvious: if it really does what it says it does, the price is an absolute bargain. That is precisely the challenge: that real experience confirms what the benchmarks say. The hardware mystery. DeepSeek has not revealed what hardware has been used to train this version of its founding model. In the past they did admit that they had used NVIDIA’s H800s. Which yes it is known The thing is that the model has been developed to run on both NVIDIA and Huawei Ascend chips. This last has confirmed Baidu that its Ascend Supernode clusters based on the Ascend 950 will fully support DeepSeek v4 versions. Huawei support is “horrible” news for the US. In The Information they already commented that one of the reasons for the “delay” in the appearance of this model was to adapt it so that it worked without problems with Huawei chips. That support is according to Jensen Huang “horrible” news for the US, because it means that dependence on NVIDIA chips no longer exists or at least is reduced to a minimum. But. The launch comes at a difficult time for the company. Guo Daya, one of the people responsible for the v1 and v3 models, has signed for ByteDance to work on AI agents. Luo Fuli, who led the development of v2, joined Xiaomi last year. This launch also coincides with DeepSeek seeking external funding for the first time. They are expected to raise about $300 million and obtain a valuation of about $20 billion. according to The Wall Street Journal. From the surprise effect to the continuity effect. The launch of DeepSeek R1 in January 2025 was surprising because it demonstrated that China could train competitive models at a fraction of the cost of Western models. With DeepSeek v4 that surprise effect disappears to give way to the continuity effect. This model seems to maintain precisely what made the previous model famous: extraordinary power at a very low cost. Bad news for Anthropic. Such low prices are terrible news for Anthropic, which in recent weeks has been forced to execute a kind of “reduflation” of their new modelswhich are not more expensive but consume many more tokens. We’ll have to see if DeepSeek v4 Pro is as good as the company promises, but if it is, we’ll have another “DeepSeek moment” before us. Maybe not as notable as last year’s, but equally relevant. In Xataka | DeepSeek promised them happiness as the great Chinese AI. I didn’t count on a small detail: Kimi

DeepSeek promised them happiness as the great Chinese AI. I didn’t count on a small detail: Kimi

Just a year ago, DeepSeek was one of the biggest scares that Silicon Valley had received dwarves. A Chinese model trained with a fraction of OpenAI’s budget equal to GPT-4 in benchmarks. Upon its arrival the message seemed clear: Western dominance of AI had its days numbered. Today, the story stands, but not thanks to DeepSeek. The DeepSeek case. DeepSeek carries months late for its V4 and, to date, has already lost three of the authors of R1, the model that catapulted them to success. The monthly downloads fell 72% in the second quarter of the year, seeing how Doubao (ByteDanec) snatched the lead. With missed dates, usage errors due to cyber attacksand the difficulty of split from NVIDIA To bet almost entirely on Huawei’s Ascend chips, Chinese alternatives like Kimi have been gaining ground. Meanwhile, on the other side of China. Moonshot AI was not born surrounded by noise like DeepSeek. It was founded in March 2023 by three former colleagues from Tsinghua University: Yang Zhilin—PhD from Carnegie Mellon, former Google Brain and Meta AI—, along with Zhou Xinyu and Wu Yuxin. There were no visible or media faces behind it, only product. That product is Kimi, and in early January 2026 the company launched it in its K2.5 version. In code and video benchmarks managed to surpass GPT-5 and Gemini Pro 3with the key to Chinese AI: its API costs between 4 and 17 times less than OpenAI’s. Those responsible for Moonshot explained how Kimi was almost at Claude’s level in software development testing, encouraging the race for open models. The money arrived. The commercial results are what really attract attention. In less than 20 days Following the launch of K2.5, Kimi’s cumulative revenue exceeded everything billed during 2025. API’s international revenue increased fourfold since November of the previous year. The consequence in valuation has been dizzying: 4.3 billion dollars in December 2025, 10 billion in February 2026, 18 billion in March. Three months, valuation multiplied by four. Kimi has thus become the fastest decacorn in Chinese business history. The Chinese maelstrom. DeepSeek was born a year ago as the great revolution that questioned the closed model of Silicon Valley. It only took a few months for Moonshot to steal the limelight and manage to be on par with – or even above – giants like Google and OpenAI in the most used models in the world. In favor of DeepSeek, it should be noted that its objective is different: it does not follow the typical startup pattern with pressure for immediate monetization and it is a gigantic AI laboratory that can afford not to win in the short term. In Xataka | DeepSeek API: what it is, what it is for, prices and how you can get one to use in your projects

Anthropic just accused DeepSeek and other Chinese companies of “distilling” Claude

For months we have talked about the race between the United States and China to dominate artificial intelligence as if it were only a question of who trains the most powerful model or launches the next version first. But the pulse begins to move to another, more delicate area: that of the rules of the game. When one laboratory accuses another of extracting capabilities from its system to accelerate its own development, the discussion goes beyond the technical. That’s exactly what Anthropic just did by denounce “distillation” campaigns against his model Claude. The complaint. In a text published this Monday, the company claims to have detected “industrial-scale campaigns” aimed at extracting Claude’s capabilities. According to their version, the activities attributed to DeepSeekMoonshot and MiniMax reportedly involved more than 16 million queries, question and answer interactions, and were channeled through approximately 24,000 fraudulent accounts, in violation of their terms of service and regional access restrictions. The race and the suspicion. The announcement by the firm led by Darío Amodei occurs in a context of growing tension around the progress of Chinese AI. Let us remember that DeepSeek altered the Silicon Valley landscape a year ago with the launch of R1, a competitive model that was presented as Developed at a fraction of the cost of American alternatives. The impact was immediate on the markets and revived the political debate in Washington about the technological advantage over China. Distilling is not always cheating. Anthropic itself recognizes that distillation is a common technique in the sector. It consists, in simple terms, of training a less capable model using the responses generated by a more powerful one, something that large laboratories use to create smaller, cheaper versions of their own systems. The problem, according to the company, appears when this practice is used to “acquire powerful capabilities from other laboratories in a fraction of the time and at a fraction of the cost” that developing them independently would entail. In that case, distillation would cease to be an internal optimization and would become, always according to Anthropic, a way of taking advantage of the work of others. Recognizable pattern. The three laboratories would have used fraudulent accounts and proxy services to access Claude on a large scale while trying to avoid detection systems. The company details infrastructures, what it calls “hydra cluster”, extensive networks of accounts that distribute traffic between its API and third-party cloud platforms, so that when one account was blocked, another took its place. Anthropic maintains that what differentiated these activities from normal use was not an isolated query, but rather the massive and coordinated repetition of requests aimed at extracting very specific capabilities from the model. Three campaigns. Although Anthropic presents the campaigns as part of the same dynamic, it distinguishes relevant nuances. DeepSeek would have focused its more than 150,000 queries on extracting reasoning capabilities and generating safe alternatives to politically sensitive questions. Moonshot, with more than 3.4 million queries, would have been oriented towards the development of agents capable of using tools and manipulating computing environments. MiniMax would concentrate the largest volume, more than 13 million queries, and according to Anthropic’s account, it reacted in a matter of hours to the launch of a new system, redirecting its traffic to try to extract capabilities from its most recent system. A geopolitical issue. The company states that illicitly distilled models may lose safeguards that seek to prevent state or non-state actors from using AI for purposes such as the development of biological weapons or disinformation campaigns. It also argues that distillation undermines export controls by allowing foreign laboratories to close the gap in other ways, while at the same time recognizing that executing these large-scale extractions requires access to advanced chips, thus reinforcing the logic of restricting their availability while, at the same time, remembering that the risk would grow if these capabilities end up being integrated into military, intelligence or surveillance systems. Images | Xataka with Nano Banana Pro In Xataka | Seedance is the greatest brutality we have seen generating video. And it has an uncomfortable message: it has surpassed Sora and Veo without NVIDIA chips

Log In

Forgot password?

Forgot password?

Enter your account data and we will send you a link to reset your password.

Your password reset link appears to be invalid or expired.

Log in

Privacy Policy

Add to Collection

No Collections

Here you'll find all collections you've created before.