๐Ÿค–
AI ๐Ÿ“… 2026-08-23 ยท 03:58 AM IST โฑ 3 min read

New Hacking Method Lets People Sneak Bad Instructions Past AI Safety Filters

Security researchers discovered a technique that hides harmful requests from AI systems until they're already processing them.

Security researchers have identified a concerning weakness in major artificial intelligence platforms, including Grok and Google's Gemini. The problem involves a method that disguises harmful instructions in encrypted code, allowing them to slip past the safety protections that these AI systems normally use to refuse dangerous requests.

Think of it like this: imagine a security guard at a building entrance who checks all visitors. This new technique is like giving someone a locked box that the guard cannot see inside. Once the person is past the guard and inside a secure room where they can open the box, the contents become visible. In this case, the "secure room" is called a trusted execution environment, which is a special protected area within a computer where sensitive operations happen.

How This Attack Works

The technique being discussed here, researchers call it "Cryptographic Context Injection," works by wrapping dangerous instructions in layers of encryption. The AI system receives what looks like harmless data and processes it normally. However, once the instructions reach a certain part of the system designed to be extra secure, the encryption gets unlocked. At that point, the harmful content can be exposed and acted upon before safety filters can catch it.

This is particularly serious because Grok and Gemini are designed with multiple layers of protection. These safeguards are supposed to prevent the AI from helping with illegal activities, creating harmful content, or producing biased results. When researchers tested this new vulnerability, they found it could potentially bypass these protections.

Why You Should Care

If you use these AI systems for work, school, or personal projects, this matters to you. These platforms are increasingly used for important tasks like writing, coding, research, and decision-making. The safety features exist to keep the AI helpful and honest.

For businesses using these AI tools, this vulnerability could expose them to legal liability and reputational damage if the systems are tricked into doing something harmful.

What You Can Do

For most everyday users, direct action options are limited, but you're not helpless:

For organizations, work with your security teams to understand the risks and implement additional verification steps for critical AI-generated content.

The companies behind these systems will likely release patches soon, but this discovery highlights an ongoing challenge: keeping increasingly sophisticated AI systems secure requires constant vigilance.

๐Ÿ“Ž This is original ITVedas reporting. This story was inspired by coverage from source. Visit the source for their original reporting.

Want to understand the technology behind this story? ITVedas has beginner-friendly guides on every IT topic.

Explore IT Chapters โ†’