Sunday, September 13, 2026

New EchoGram Trick Makes AI Models Accept Dangerous Inputs

Security researchers at HiddenLayer have uncovered a critical vulnerability that exposes fundamental weaknesses in the guardrails protecting today’s most powerful artificial intelligence models.

The newly discovered EchoGram attack technique demonstrates how defensive systems safeguarding AI giants like GPT-4, Claude, and Gemini can be systematically manipulated to either approve malicious content or generate false security alerts.

How the Attack Works

The EchoGram technique exploits a shared vulnerability across the two most common AI defense mechanisms: classification models and LLM-as-a-judge systems.

Both approaches rely on curated training datasets to distinguish between safe and malicious prompts.

EchoGram targets Guardrails, unlike Prompt Injection which targets LLMs
EchoGram targets Guardrails, unlike Prompt Injection which targets LLMs

By identifying specific token sequences underrepresented in these training datasets, attackers can “flip” the verdicts of defensive models, causing them to misclassify harmful requests as benign.

What makes EchoGram particularly dangerous is its simplicity. A researcher testing an internal classification model discovered that appending the string “=coffee” to a prompt during a prompt-injection attack caused the guardrail to approve the malicious content incorrectly.

This seemingly random string represents a calculated exploit that exploits imbalanced training data.

The attack operates in two troubling ways. First, attackers can append nonsensical token sequences to malicious prompts, bypassing security filters.

At the same time, the harmful instruction still reaches the underlying language model intact. Second, researchers demonstrated that EchoGram can generate false positives by crafting benign queries containing specific token combinations.

This could flood security teams with incorrect alerts, making it harder to identify genuine threats.

HiddenLayer’s testing revealed that a single EchoGram token successfully flipped verdicts across multiple malicious prompts in commercial models.

Even more concerning, combining multiple EchoGram tokens created powerful bypass sequences that degraded a model’s ability to identify harmful queries.

Tests on Qwen3Guard, an open-source harm classification model, showed that token combinations could flip safety verdicts even across different model sizes, suggesting a fundamental training flaw rather than an isolated issue.

EchoGram flipping the verdict of various prompts in a commercially available proprietary model
EchoGram flipping the verdict of various prompts in a commercially available proprietary model

The research highlights a critical problem in the ecosystem. Many leading AI systems use similarly trained defensive models, meaning an attacker who discovers one successful EchoGram sequence could reuse it across multiple platforms, from enterprise chatbots to government AI deployments.

This vulnerability isn’t isolated; it’s inherent to current training methodologies.

The discovery exposes a false sense of security that has developed around AI guardrails. Organizations deploying language models often assume they’re protected by default, potentially overlooking deeper risks.

Meanwhile, attackers can exploit this misplaced confidence to either slip past defenses undetected or undermine security team confidence through alert fatigue.

Benign queries + EchoGram creating false positive verdicts
Benign queries + EchoGram creating false favorable verdicts

EchoGram represents a wake-up call for the AI safety community. As language models become embedded in critical infrastructure across finance, healthcare, and national security, their defenses require continuous testing, adaptive mechanisms, and transparency in training methodologies.

HiddenLayer emphasizes that trust in AI safety tools must be earned through demonstrated resilience, not assumed through reputation alone.

The research underscores an urgent need for the industry to move beyond static defenses toward dynamic systems capable of withstanding emerging attack vectors.

Follow us on Google NewsLinkedIn, and X to Get Instant Updates and set GBH as a Preferred Source in Google.

Divya
Divya
Divya is a Senior Journalist at GBhackers covering Cyber Attacks, Threats, Breaches, Vulnerabilities and other happenings in the cyber world.

Hot this week

How To Access Dark Web Anonymously and know its Secretive and Mysterious Activities

What is Deep Web The deep web, invisible web, or...

How to Build and Run a Security Operations Center (SOC Guide) – 2023

Today’s Cyber security operations center (CSOC) should have everything...

Russian Hackers Bypass EDR to Deliver a Weaponized TeamViewer Component

TeamViewer's popularity and remote access capabilities make it an...

Web Server Penetration Testing Checklist – 2026

Web server pentesting is performed under three significant categories: identity,...

ATM Penetration Testing – Advanced Testing Methods to Find The Vulnerabilities

ATM Penetration testing, Hackers have found different approaches to...

Threat Actors Use Claude AI Agents to Automate Cyberattacks and Steal Sensitive Data

Threat actors are increasingly using Claude-based AI workflows to...

China-Linked Hackers Chain Chrome Zero-Day With Windows Kernel Flaw in Attacks

China-linked threat actors UTA0560 and JungleBamboo chained a Google...

New Phishing Campaign Abuses Windows Mshta.exe to Steal Credentials and Secrets

A newly identified phishing campaign is abusing the legitimate...

CISA Warns of Critical GitLab Vulnerability Exploited in Attacks

The U.S. Cybersecurity and Infrastructure Security Agency (CISA) has...

Researchers Uncover 10,000+ Malware Loaders Behind YouTube and SEO Poisoning Campaign

A long-running pay-per-install (PPI) operation that used YouTube gaming...

VLC Media Player Flaws Let Attackers Corrupt Memory and Leak Sensitive Data

Two security vulnerabilities in VLC media player versions 3.0.0...

CISA Adds Exploited MikroTik RouterOS Flaws to Security Alert

The U.S. Cybersecurity and Infrastructure Security Agency (CISA) has...

Related Articles

Recent News