7 Deadly Signs of AI Security Snake Oil: A Developer's Field Guide

Seven signs you're being sold snake oil when it comes to AI vendors and selling framework coverage as "protection"

Gray Swan
August 17, 2026
“Cookery simulates the disguise of medicine... it pretends to know what foods are best for the body.” - Plato 

Over 2000 years later and we still find ourselves enmeshed in various iterations of Plato’s sentiment. AI Security not withstanding.

AI vendors love to scare you. Conjure up a dash of panic with a pinch of shock. Then they produce the ‘antidote.’


What’s in the vial?

A very special mixture: their OWASP Top 10 coverage and their MITRE ATLAS compliance. ‘It’s comprehensive. Aligns with any framework.’

Comprehensive? Any framework? The confidence is alluring. Your shaky hand considers the dotted line.  

Wait. 

A glaring problem. None of those ingredients indicate whether or not YOUR deployment is secure. Not a drop of consideration has been given to your system’s particulars. A lot of what passes for ‘AI Security’ is nothing more than second-rate cookery pawned off as an all-encompassing medicine for security threats. 

But how can you tell the difference?

There are 7 Deadly Signs that a vendor selling “peace of mind” shouldn’t even exist. Let’s explore them and see if there’s not some kind of alternative. An alternative where there isn’t a drop of snake oil in sight. 

Sign #1: They talk about frameworks more than your system. 

“Let’s talk about frameworks, baby. We cover all OWASP categories. You name it, we cover it.”

The potential vendor leans back in his chair, flashing the pearliest set you’ve ever seen. He speaks in platitudes. Generalities. If we could, we’d stand behind him waving a battle-sized, red flag.
Their mistake? At the end of the day, it’s not just about OWASP categories. That means very little for your own particular system. 

“How about prompt injection?” he says. 

What about it? They still haven't asked anything about your own particulars. Prompt injection means different things. Do you have a coding agent with terminal access? Or a customer service chatbot? One can delete your codebase during normal operations. The other might leak customer data through a malicious email. 

The first question AI security vendors ask you shouldn’t be “What framework do you need to comply with?”

It ought to be, “Tell me what your agent can actually do.”  “What tools does it have access to? What data can it see?” 

Why so specific? Because that’s what determines your real attack surface. MITRE ATLAS covers over a hundred different attack techniques across different AI platforms. You can’t expect arthritis relief from a pill promising to alleviate stomach pains.

Your particular threats runneth over. And their promises? Empty. 

Sign #2: They source threat intelligence from the models they're protecting.

In football there’s a popular torture method coaches deploy commonly referred to as the “Oklahoma Drill.” Two players face each other in proper stance and plow into each other at the whistle. 

Now imagine a coach saying, “We’ve proven that our blocker can stop any lineman.” 

Big if true, but you have to ask: “Whom did you test him against?”

They straighten their posture and try to stymie the proud grin on their face. “We used our own nose tackle.” 

When purveyors of so-called AI security use their own model to generate attacks against itself, they’re verging on circular reasoning. If they’re gleaning data using this method, that means both their attacker agent and their model have shared blind spots. They’ve created an echo-chamber.

AI cyber attacks vary widely. And a company’s own system is uniquely vulnerable. A financial enterprise has different needs from a company dealing in biomedical research. If a vendor has been using GPT-4 to generate attacks against GPT-4, their data is both sanitized and outdated.

If the nose- tackle always cuts left, you can’t expect the blocker to anticipate a cut to the right. 

See the problem? An honest and competent vendor will have robust data. Will understand the ins and outs of AI security based on the experience of boots-on-the-ground red-teamers. They’ve taken a both creative and empirical approach to AI Security. 

And they continually glean new data. When a novel threat presents itself on a Tuesday, they have it recognized and documented by Wednesday. Then it’s time for a real time solution. 

Sign #3: They’re talking supply chain when you’re using frontier models. 

You’ve just purchased a house in a place with high rainfall. A homeowner’s insurance agent keeps pushing protection from wildfires. Not a peep about floods.

Vendors like to glance at their MITRE ATLAS cheat sheet and throw out words like ‘model poisoning’ and ‘supply chain attacks.’ 

The problem? Those don’t apply to a company utilizing an API provider. You never touch the training pipeline or see the infrastructure. Essentially, the vendor is testing for threats that ended the moment you selected an API provider. 

They’re testing for fortitude against fires rather than the impending floods. 

Ask yourself this question: “Which half of MITRE ATLAS actually applies to my deployment?” 

Chances are half of it doesn’t. And it’s that half vendors like to push. This shows a fundamental ignorance of your particular system. 

The only way to mitigate ignorance like this is testing within a replica of your own system’s environment. A quality vendor understands that ‘prompt injection’ and ‘data poisoning’ are some of the critical issues facing a company with an API provider. 

Don’t get burned by misapplied MITRE ATLAS terminology. They’re selling you asbestos solutions when you ought to be building dams. 

Sign #4: They claim ‘model-level safety’ solves your problem. 

You’ve hired an engineer to build the strongest wall around your city. The problem? Nobody gave a thought to the sewers. When the enemy arrives, the wall’s strength will be irrelevant. 

Even the safest model in the world doesn’t know PII tables or financial data are HIGHLY sensitive. It isn’t aware that ‘employee_payroll’ and ‘bank_accounts’ ought to be stamped with a big, red TOP SECRET. 

The safest model doesn’t know which API calls cost you money and time. It doesn’t know that ‘send-email’ shouldn’t be used to exfiltrate sensitive data. 

That’s system-level knowledge. 

Your city commissioner has the map of the sewers. The wall engineer doesn’t have the faintest idea about them.

Defenses NEED to operate at the runtime layer. It’s where your agent meets your tools, data, and policies. 

Identical models can exhibit completely different behaviors, depending upon their access and permissions. It’s all about what YOU let them do. 

About YOUR vulnerabilities. 

Sign #5: They can’t explain what happens when your agent goes off-policy without an attacker.

A well-defended city doesn’t always need an outside enemy to push it toward collapse. Sometimes the enemy is within.

For instance, you’ve prompted your agent to delete “old files.” After the fact, you realize all recovery files have been eliminated from the system. Cue hours or even days of mayhem. That’s time and money.

It’s not always a jailbreak that wipes out a codebase. Sometimes it comes down to prompt specificity and missing context. An agent might do something drastic, all while still being within the technical bounds of its specific policies and parameters. A simple misunderstanding of context can lead to headaches and tragedies for a company utilizing various agents.

A customer service agent doesn’t need an adversary to leak internal information. Its breach of conduct could be an attempt to optimize for helpfulness.

If your potential vendor only tests for adversarial attacks, they’ve potentially missed half the failure modes. Even during normal operation, your agent can violate policy.

Don’t bother with “Can somebody jailbreak this?” The real question? “Does my system behave according to its intended specification?”

It may very well be a mole within your own walls. And any vendor worth their salt is going to acknowledge and plan for just that. 

Sign #6: They focus on generic attacks instead of deployment specific behavior.

Vendors who start with “Can someone jailbreak this?” will often show you correlating alerts.

WE DETECTED AN SQL INJECTION ATTEMPT! 

Again, that is neither the right question nor the right test. At all times, when dealing with AI security, the question addresses system particulars. 

“Does this system behave according to MY requirements?” Sometimes a nefarious party will use indirect prompt injection through a malicious customer email to make your service agent leak PII. 

Other times they’ll manipulate your coding agent. Hit them with a hidden README to introduce a back door. 

These two examples are fundamentally different problems:

  • One has to do with general robustness in the face of known attack patterns. 
  • The other is about enforcement. Does your system enforce your policies with your tools in your environment? 

Most vendors are only concerned with the first example. They speak only in terms of what they don’t understand. What don’t they understand? The specifics of your system. 

A cat burglar pillaging a street will consider the particular vulnerabilities of each individual house he plans to enter. He already assumes the doors will be locked. Some vendors will test if the doors are locked, and then tell you the house is protected. 

A vendor worth their salt won’t run static benchmarks. They specify attacks based upon your particular deployment. Prompt injection’s existence isn’t an automatic vulnerability. Specific questions are essential. And they happen to be our bread and butter. 

Can someone exploit YOUR agent's access to YOUR GitHub and YOUR Slack to do something YOUR policy prohibits?"

Sign #7: They describe AI security as easy or simple ‘Cybersecurity’.

As it turns out, terminology matters. Let’s break down some of the semantics. This one is at the bottom, because it’s a foundational and deplorable sign. When it comes to AI security, most traditional cybersecurity rules become secondary concerns. That is to say, semantics are important.

AI Security vs. Cybersecurity

There’s an ever clear and present distinction to be made between ‘AI Security’ and ‘Cybersecurity.’ The ways and means of protecting your system are changing rapidly. Few have a firm grip on those changes.  Even fewer have a grip on how to harness them. 

Securing agentic AI is inherently contextual. The right defenses for a coding agent are totally wrong for a financial trading agent. When a vendor waves it off as ‘easy,’ they are telling on themselves. Applying an understanding of traditional Cybersecurity measures to AI Security is a gut wrenching boondoggle.

A gross oversimplification that verges on criminality. 

Why do they think it's easy? Because they spent so much time acquiring the vocabulary. They neglected the part where they needed to understand the application of those terms. 

Stochastic vs. Deterministic

A vendor perhaps understands that traditionally, a system’s inputs and outputs can be kept consistent. These traditional systems are deterministic. An input will produce a consistent output. But AI systems are stochastic.  They generate varying outputs for the same input. Your specific system needs an entire list of unwanted outputs. A rigid framework to ensure it doesn’t make a mistake while still remaining on script. 

To call ‘AI Security’ simple ‘Cybersecurity’ is completely missing the point. 

AI and its security are in the Wild West phase of their development. It’s hard for anyone to understand the ins and outs of the technology. The implications for productivity and the particular associated risks. Anyone smarter than the average bear can see nothing about it is easy. And when that vendor flashes their pearly grin and waves it all off as ‘simple?’ Or worse, ‘simple cybersecurity’?

Return the grin and wave them off. They don’t  know what they’re talking about. 

You’re not mechanics discussing an oil change. 

Signs and their solutions: 

In AI Security, a lot of vendors miss the mark. They talk a big game. Talk about frameworks and supply chain. They make frivolous claims about ‘model-level safety.’ But they can’t explain what happens when your agent goes off-policy without an attacker. And any discussion of attacks concerns the generic variety. They ignore system particulars in favor of ‘model-level safety’ solving all of your problems. 

And at the core, when it comes to AI Security, they just don’t get it. They’re mentality is stuck somewhere back in the 2010s. 

It’s easy to point out problems. Identify the flaws. It’s all well and good to understand the problems, but what about solutions? We’ve identified what these snake oil vendors could be doing right, but what does that look like? 

Humor us as we self-reflect. Make a few observations about concrete solutions with which we’re familiar. A sample, if you will. 

  1. Sourcing Threat Intelligence from Models They’re Protecting

Antidote: Gray Swan hosts The Arena. This is a global platform where researchers test AI models’ vulnerabilities. Over 15,000 red-teamers attempt to break the latest models. And they haven’t failed yet. We harvest ample data from this enterprise, and transfer the knowledge gleaned from The Arena data to those searching for AI Security solutions. Our understanding of AI Security comes directly from the boots on the ground. Practices are driven by thorough and diverse sets of data. 

  1. “Model-level safety solves your problems” 

Antidote: We operate our defenses at the runtime layer. Create a replica of your total environment. That way all tests can be done where your agent meets YOUR tools, YOUR data, and YOUR policies. Identical models can exhibit completely different risk profiles. It all depends upon what you LET them do. And we understand that. Our primary concern is always figuring out the specific particulars of your individual and distinct system. 

  1. Focusing on generic attacks instead of deployment-specific behavior. 

Antidote: Our attack agent, Shade, doesn’t run static benchmarks. Rather, it generates attacks based upon your own specific deployment. It works with your own tools, data access patterns, and policy constraints. We ask questions about how a potential attacker could exploit your agent’s access to your particular system. The priority is system specifics rather than vague attack categories. 

A knowledgeable vendor needs to have particular solutions. They should be able to both point out the problems plaguing the AI Security space and offer solutions for them. Those solutions should be driven by two factors: thorough data and your system’s particulars. 

Spotting the signs is one step. Addressing them is a different matter entirely. 

Conclusion:

Frameworks give you vocabulary, NOT security. Your system’s risk profile is determined by what it can actually do. It has nothing to do with checking taxonomy boxes.

Take this field guide checklist on your next vendor call. You’ll find something disquieting.

The veneer of cutting edge technology makes it difficult to assess a pitch’s validity. As does the speaker’s confidence. But when you take a step back and ask these questions, you’ll realize something. A hundred fifty years ago, the peddler wheeling into the next town deployed the same tactics. A magical tincture for any and all aches, pains, and inner ailments. Dr. Fantastico’s Finest Liniment of Snake. 

Buyer beware.