Deepseek does the same as Openai’s most advanced models with much less resources. The key: “Reinforcement Learning”

The entire world is wondering how it is possible that the models of AI of Deepseek They have become overnight the great protagonists of today in the field of artificial intelligence. The answer is relatively simple. These models have managed to demonstrate that You can do more with much less. Both Deepseek V3 and Deepseek-R1 are comparable to GPT-4 or O1 OPENAI respectively, but it is estimated that their training has been much less expensive and its inference, of course, is: the prices of the Deepseek API are up to 35 sometimes lower than those of OpenAi, but that makes one wonder how it is possible. The answer is clear, and it is because we have at our disposal the technical reports of these AI models. Precisely his study has allowed us to clarify What are the techniques that this Chinese R&D laboratory has used to develop these models so efficient and capable. Many techniques, a single objective: efficiency There are several differences that make Deepseek’s new model especially efficient. Its creators explain in detail in the detailed Technical Report that is publicly available. Here are the most relevant: Deepseekmoe (“Mixture of experts”): In models such as GPT-3.5 the entire model was activated in both training and inference (when we use it). However, not all model components are necessary for our requests. The MOE technique – already introving with Deepseek V2 – precisely divides the model into multiple “experts” and only activates those that are necessary according to the request. GPT-4 is already a MOE model. But as we said, Depseekmoe even went further and differentiated between even more specialized experts, in addition to using some somewhat more generalist experts that could contribute value in certain requests. Managing all those specialized or generalist experts not only benefits inference, but also the training phase, making it more efficient. This technique is similar to the so -called “Time Scaling test” that also adjusts the size or complexity of a model during efficiency. Deepseekmla (Multi-Head Latent attention): It is another substantial improvement-even more than the previous one, and also introduced with Deepseek V2-that affects the way in which memory is managed in these models. Normally it is necessary to load both the model and the entire context window – the one that allows us to write prompts and include long texts, for example. Context windows are especially expensive because each token requires both a key and their corresponding value. With the improvement introduced with this technique, what was made possible was to compress that warehouse of keys and values, dramatically reducing memory use during inference. Auxiliary -los-Free Load Balancing: If we imagine a model like a great orchestra, each musician is an “expert” within the model. To play a complex piece, not all musicians are necessary all the time. Traditionally the so -called “auxiliary losses” were used to make sure that all musicians played enough, but these losses could interfere with that interpretation of the musical piece (model training), which could degrade general performance. With Deepseek V3 the model is able to balance the work of each expert dynamically. That does the simplest, direct and efficient training by eliminating “auxiliary losses.” In addition, the elimination of interference allows the model to learn better and with less resources … and get better results. Multi-Token Prediction Training Objective: Often predicting the following word depends on several previous words or context. With this technique instead of predicting only the following word, the model learns to predict several words at the same time. That makes more natural and understandable and less ambiguous texts generate, but also accelerates training by reducing the number of steps necessary to generate the complete text sequence. FP8 Mixed Precision Training: The use of Numbers FP8 allows significantly reducing memory consumption and accelerates calculations. Some critical parts of the model continue to use FP32 training to guarantee precision, but there is another additional benefit of FP8: the size of the models is reduced. Other models use techniques such as quantization or parameter pruning. Although Openai does not give data on GPT-4 in this section, the assumption is that it works with BF16, more expensive in terms of memory. Although FP8 theoretically leads to less precise models, other complementary techniques such as fine-grained quantization are used to reduce the negative impact of values ​​that come out of the common, which makes a stable training possible. Cross-Node All-to-Lall Communication: During training it is necessary to constantly exchange information between all nodes (computers) connected in training data centers. That can become a bottleneck, but these new Deepseek V3 techniques include efficient communication protocols, data traffic reduction and efficient synchronization to accelerate training and, once again, reduce the costs of that process. Reinforcement and “distillation” learning as keys But in addition to all these techniques, those responsible for Deepseek V3 explain how they pressed it with 14.8 billion tokens, a process to which a supervised adjustment followed (Superved Fine-Tuning, SFT) and several stages of Reinforcement Learning (Reinforcement Learning, RL). The SFT phase-which is mentioned in the Deepseek V3 report-was completely omitted in the case of Deepseek-R1. However, learning by reinforcement is an absolute protagonist in the development of both models, especially in R1. The technique is well known in the field of artificial intelligence, and it is as if we trained a dog with prizes and punishments. The model learns to respond better by giving rewards if you do well. Over time, the model learns to take actions that maximize long -term reward. In Deepseek, learning for reinforcement is used to break down complex problems in smaller steps. In it Deepseek R1 technical report It also indicates how this model makes use of RL techniques directly on the base model, without the need for supervised training. That saves computing resources. The call also comes into play here Thought chain (chain-of-though)also mentioned in the technical report. This refers to the ability of a language model to show the intermediate steps of its reasoning. The model not only … Read more

Octopus tentacles have their own “brain.” We are now learning the implications

Octopuses are invertebrate animals, but the absence of a central nervous system like that of birds or mammals does not make their brains less interesting than the rest. Brains, emphasizing the plural since neuronal systems of each of its extremities They have a degree of independence, which leads many to consider them as such. A nervous system not at all central. Now, a group of researchers has studied the nervous systems of these cephalopods to better understand how these nine neural organs operate together and to what extent they maintain their independence. What they observed is that each of these brains had the ability to operate individually. The team responsible for the new study believes that it is thanks to the unique segmentation of the nervous system of octopuses that these animals achieve the level of skill in the management of extremely flexible organs that serve these animals to move, feed, sense their environment, and even copulate. “If you are going to have a nervous system that is going to control such dynamic movement, that is a good way to organize it,” explained in a press release Clifton Ragsdale, co-author of the study. “We think it’s a feature that evolved specifically in soft-bodied cephalopods with suction cups to carry out these worm-like movements.” Studying segmentation. The new study focused on segmentation of this curious neuronal system, analyzing the distribution and function of the neurons of these tentacles, taking as reference an octopus of the species Octopus bimaculatus. Neurons that together add up to a greater number than the neurons located in the “central brain” of the animal, which is responsible for coordinating actions that require the use of various arms. These neurons in the extremities are concentrated, explains the teaminto an axial nerve cord (ANC), which “snakes” the tentacle connected to each of the animal’s suction cups. Neural columns. The ANC analysis showed that neurons in the octopus’s limbs were grouped into “columns” that in turn formed segments that the team compared to corrugated pipes. The segments were in turn separated by gaps called “septa” from which nerves and blood vessels made their way to the muscles of the limb. “From a modeling perspective, the best way to organize a control system for this long and flexible arm would be to divide it into segments,” Cassady Olson added.co-author of the study. “There must be some kind of communication between the segments, which you can imagine attenuates their movements.” Job details can be found in an article published in the magazine Nature Communications. Much to investigate. The tentacles of octopuses are very versatile limbs that allow this animal to navigate the seabed, but also, through their suction cups, they allow these octopods to perceive the world around them, hunt and feed on their prey. Knowing the details of the functioning of such complex limbs will still require new research. In Xataka | Octopuses are not aliens, and scientists have had to come out to explain why Image | Theasereje, CC BY-SA 4.0

Log In

Forgot password?

Forgot password?

Enter your account data and we will send you a link to reset your password.

Your password reset link appears to be invalid or expired.

Log in

Privacy Policy

Add to Collection

No Collections

Here you'll find all collections you've created before.