During the last two years, he prompt injection It has been the favorite weapon of cyber attackers against artificial intelligence systems. This method involves hiding a malicious instruction within an email, a calendar invitation or a web page so that an AI agent ends up obeying the intruder instead of its legitimate user. The good thing is that this same technique also works to stop attackers.
What has happened? Cybersecurity firm Tracebit has published a research in which it demonstrates that an AI agent that is attacking a system can be stopped by simply placing, next to the passwords and keys that the attacker seeks to steal, a text designed to activate the model’s own rejection mechanisms.
As soon as the attacking AI reads that text, its security barriers go up and the attack stops in its tracks. Researchers have dubbed the technique context bombing.
How it works. The logic is simple to understand although its discovery was not so simple. All major language models have certain issues built into them that they refuse to address under any circumstances, whether for security reasons or for political and regulatory reasons. Tracebit takes advantage of just that by placing near the fake credentials (the same ones they use as bait to detect intruders) fragments of text that push the model into those forbidden areas.
According to explained Andy Smith, co-founder and CEO of Tracebit told Ars Technica, with this technique a rejection mechanism is being activated in the context, something that has a strong and forceful effect and from which agents have a hard time recovering. Once that content enters its context, the model continues to reject any subsequent instructions.
The examples, depending on the model. Not all models stop at the same thing, so the trap must be adapted to each case. For Western models, such as those from Anthropic or Google, it has worked to include references to the manufacture of biological weapons, such as instructions for producing inhalable anthrax spores.
For the models developed in China, however, the key has been to mention “Tank Man”, the man who stood in front of the tanks in the Tiananmen massacre of 1989, a topic that these systems are prohibited from boarding due to censorship by the Chinese Government and that from Xataka we have actually confirmed every time we have had to test a Chinese language model.
The results, in figures. Tracebit tried the technique about five of the most advanced models of the moment (Opus 4.8, Gemini 3.1 Pro, GLM 5.2, DeepSeek 4 Pro and Kimi 2.6) simulating an Amazon Web Services (AWS) environment and entrusting them with routine development tasks that, along the way, led them to come across false keys. In 152 recorded attack attempts, the investigation point that:
- Attacks that gained administrator access fell from 57% to 5%.
- Those who also managed to leave a permanent back door went from 36% to 1%.
- The percentage of attempts that completed at least some phase of the attack dropped from 91% to 15%.
The most striking case was that of Opus 4.8, the most effective model of the five when it came to attacking, since without the context trap it achieved administrator access in 93% of the attempts; with it, it failed in 100% of the cases.
Vulnerability by nature. The beauty of this technique is that it exploits a weakness that is very difficult to correct. And according to they count According to researchers, a developer can retrain a model to be less strict when faced with a security question, but it is much more difficult to eliminate the political or regulatory restrictions that come as standard, because they are not a failure that can be “fixed”, but rather a deliberate design decision. That makes these barriers a fairly stable point of support for those who are dedicated to developing protections against these attacks.
Background. In May, Tracebit already had presented a “canary” systemthat is, fake cloud resources that look legitimate but are not actually used, so that any attempt to access them triggers an alert for the security team. In that first test, with 951 attack attempts, the canaries gave an average of eight minutes’ notice before the attacker gained administrator access.
The problem is that the AI agents needed only 14 minutes on average to complete the attack, so the defenders had a window of just six minutes to react. Context bombing was born precisely to gain that missing time, stopping the attack instead of simply warning of it.
What this does not solve. Neither Tracebit nor anyone else claims that this is the definitive solution to prompt injection. There is still no way to prevent a model from occasionally confusing an instruction with any piece of data within its context, and that confusion is precisely what makes the technique possible.
Cover image | AI-generated with Gemini
In Xataka | “It could weaken US competitiveness”: Kimi K3 forces Trump to resume his plan to stop Chinese AI


GIPHY App Key not set. Please check settings