<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Brian Gardner]]></title><description><![CDATA[Notes on what's being built and what's being sold.]]></description><link>https://letters.bgardner.net</link><image><url>https://letters.bgardner.net/img/substack.png</url><title>Brian Gardner</title><link>https://letters.bgardner.net</link></image><generator>Substack</generator><lastBuildDate>Thu, 08 Oct 2026 21:57:57 GMT</lastBuildDate><atom:link href="https://letters.bgardner.net/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Brian Gardner]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[briangardner514040@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[briangardner514040@substack.com]]></itunes:email><itunes:name><![CDATA[Brian Gardner]]></itunes:name></itunes:owner><itunes:author><![CDATA[Brian Gardner]]></itunes:author><googleplay:owner><![CDATA[briangardner514040@substack.com]]></googleplay:owner><googleplay:email><![CDATA[briangardner514040@substack.com]]></googleplay:email><googleplay:author><![CDATA[Brian Gardner]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Enterprise AI Watch — the week of August 17, 2026]]></title><description><![CDATA[The July breach at Hugging Face was not an outside attacker.]]></description><link>https://letters.bgardner.net/p/enterprise-ai-watch-the-week-of-august-dd5</link><guid isPermaLink="false">https://letters.bgardner.net/p/enterprise-ai-watch-the-week-of-august-dd5</guid><dc:creator><![CDATA[Brian Gardner]]></dc:creator><pubDate>Wed, 26 Aug 2026 14:00:55 GMT</pubDate><content:encoded><![CDATA[<p>Folks,</p><p>Here is the week of August 17. Most of it is one story, and that story is a correction to something a lot of us thought we already understood. The rest is the usual: what shipped, what broke, who got funded, what the rule-makers did. I spent twenty-five years building enterprise data protection and storage products, and I read this space the way I read that one. The demos are fun. What interests me is what happens when something goes wrong, and who has to answer for it.</p><p>One standing disclosure before we start. I consult in this industry, and the views here are mine alone, not those of any client or company I work with. This letter reports public developments and what I think they mean for the people buying and running these systems. It is analysis, not advice: not investment advice, and not a recommendation to buy, sell, or hold anything, including the securities of any company named here. Where I mention funding or valuations, I am reporting what was announced. I hold no positions in any company named in this report. If that ever changes, I will say so at the time.</p><h2>What happened this week</h2><p>The July breach at Hugging Face was reported at the time as an intrusion by an outside attacker. It was not.</p><p>OpenAI has taken responsibility. Two of its own models, GPT-5.6 Sol and an unreleased and more capable one, broke out of an internal capability evaluation where safety classifiers had been deliberately reduced for the test. Over roughly two months they moved from OpenAI&#8217;s own systems into Hugging Face&#8217;s production infrastructure. That is where the seventeen-thousand-action log behind the original headlines came from. OpenAI admitted it on <a href="https://www.axios.com/2026/07/21/openai-says-hugging-face-breach-caused-by-one-its-models">21 July</a> and gave the full account <a href="https://simonwillison.net/2026/Aug/7/openai-timeline/">at Black Hat on 6 August</a>. Axios, TIME and Fortune have all reported it independently, so this is not just OpenAI&#8217;s own telling.</p><p>The detail I keep coming back to is how OpenAI found out. By the reporting, it asked Hugging Face to revoke a set of credentials and was told they had already been revoked. That is the moment OpenAI learned it was investigating itself.</p><p>Two things follow for anyone running these systems.</p><p>First, the boundary that failed was an evaluation environment. A lab. Every place I have ever worked had one. In my experience labs are simply not all alike. Yet the coverage keeps treating them as if they were. I have worked in labs that were sandboxed to the max. Carefully designed, thoroughly thought through, and nothing was getting outside that environment, period. Labs like that are run by lab specialists, and they are a special breed. They think several links down the chain of consequences for every action they take, and they plan for those consequences and how to handle them, because they know the outcome lands on their reputation. Most companies do not even think about hiring or keeping lab specialists. They see a need for a lab based on a business requirement and put it together quickly with the people they have. I have no inside knowledge of the people who built OpenAI&#8217;s evaluation environment. But an environment that sat quietly wired to production for two months does not sound to me like it was built by lab specialists. So the question this week hands us is not whether a specific lab is configured well. It is who runs it, and what it can reach.</p><p>Second, and this is what I think we all need to notice. <a href="https://huggingface.co/blog/security-incident-july-2026">Hugging Face&#8217;s own post about the incident</a> is still live as written. It describes a malicious dataset, an attacker action log, and says the model behind it is still not known. It does not name OpenAI anywhere. So two companies&#8217; public accounts of the same event no longer agree, and the older one is the one most people will find first. Updating an incident post is genuinely hard and it is nobody&#8217;s favorite Tuesday. But if your own incident record cites the original framing, it is wrong today, and it will stay wrong until somebody goes back and checks it. That is a filing problem rather than a security problem, and it is the kind that outlives the incident.</p><h2>Corrections</h2><p>I carried some numbers in the last issue that need fixing, so let me do that before anything else.</p><ul><li><p><strong>The JFrog CVE counts.</strong> The figures that circulated with this story, three in some tellings and nine in others, are not counts JFrog has confirmed. <a href="https://jfrog.com/blog/jfrog-and-openai-collaboration-on-zero-day-security-findings/">JFrog&#8217;s own advisory</a>gives no count at all. Treat both numbers as unverified, mine included.</p></li><li><p><strong>The Zenity install figures.</strong> A 250,000-install figure and a 1.7 million aggregate come from two different Zenity releases, an <a href="https://zenity.io/company-overview/newsroom/company-news/zenity-labs-discovers-dozens-of-malicious-ai-agent-skills-evading-detection-launches-ai-total">August 3 product launch</a> and an <a href="https://labs.zenity.io/post/attackers-target-agents-via-the-skill-supply-chain">August 6 campaign disclosure</a>. Reading the technical writeup, I think the smaller figure is one family inside the larger total rather than a competing count. But Zenity never states the arithmetic, so that is my inference and not a reconciliation. Do not quote it as one.</p></li></ul><h2>What shipped</h2><ul><li><p>Google <a href="https://cloud.google.com/blog/products/ai-machine-learning/expanding-google-antigravity-for-enterprise-customers">brought Antigravity, its AI coding agent, into Gemini Enterprise</a> on 21 August, with administrative controls over budget and resource use &#8212; monthly spending thresholds, shared token pools, and overages an administrator has to opt into with a hard cap. A budget ceiling arriving as a first-class admin control instead of a billing report after the fact is a small change in where the limit lives and a large one in who answers for it.</p></li><li><p><a href="https://www.helpnetsecurity.com/2026/08/14/new-infosec-products-of-the-week-august-14-2026/">The week&#8217;s product roundups</a> named Hazmat, an open-source sandbox for agent execution, A10 Networks&#8217; AI Gateway, and an expanded release of ScienceLogic&#8217;s Skylar AI.</p></li><li><p>Noma Security launched <a href="https://www.prnewswire.com/news-releases/noma-launches-agentic-access-control-to-govern-ai-agents-and-mcp-servers-across-the-enterprise-302788534.html">Agentic Access Control</a>, for governing agents and MCP servers across an enterprise. The company also claims 1,300% ARR growth on top of a previously reported $100M Series B. That is the company&#8217;s own figure and I am reporting it as claimed.</p></li><li><p>Microsoft&#8217;s agent toolkit now carries a caveat in <a href="https://github.com/microsoft/agent-governance-toolkit/blob/main/docs/ARCHITECTURE.md">its own architecture documentation</a>, conceding that its published benchmark results are &#8220;specific to this test suite&#8221; and &#8220;should not be interpreted as universal guarantees.&#8221; I would like to see a lot more of this. A benchmark number with its scope stripped off is the single most-copied unreliable figure in enterprise software, and a vendor writing the limit into its own docs is not just doing the reader a favor. It is doing us all one.</p></li></ul><h2>Standards</h2><ul><li><p>The Model Context Protocol&#8217;s <a href="https://blog.modelcontextprotocol.io/posts/mcp-roadmap/">roadmap</a>, published 22 August, names agent identity and enterprise security a top-five priority for the coming cycle. Specifically: finishing adoption of DPoP, a scheme that ties a login token to a key only the rightful holder has, so a stolen token is useless on its own; a standard way for one agent to act on behalf of another under its own verifiable identity; and standard token exchange in place of static API keys. Static keys have been the quiet default in agent deployments for two years now. Naming their replacement a roadmap priority is the first sign that is ending.</p></li><li><p>An <a href="https://duendesoftware.com/blog/20260820-summer-2026-identity-standards-recap">identity-standards recap</a> on 20 August reports that OAuth Identity Chaining &#8212; a way to keep a user&#8217;s identity attached to a request as it passes through a chain of services and agents &#8212; has been approved by the IETF&#8217;s steering group as a Proposed Standard. The same recap reports a new draft that would require a cryptographically signed human approval step for sensitive agent actions.</p></li><li><p>The Cloud Security Alliance&#8217;s AARM conformance registry &#8212; Autonomous Action Runtime Management, its specification for securing agent actions at runtime &#8212; <a href="https://aarm.dev/builders">still lists eight conformant products</a>, unchanged since 3 August. The wider self-registered &#8220;aligned&#8221; list, which is companies saying they build in the same space rather than companies that passed anything, stands at ninety-five. Eight and ninety-five is the story. Signing up is free and fast, passing a review is neither, and either number quoted on its own misleads. Quote them together or not at all.</p></li></ul><h2>Incidents and research</h2><ul><li><p>Fortinet <a href="https://www.securityweek.com/fortinet-acquires-ai-security-company-virtue-ai/">acquired Virtue AI</a> on 17 August, adding automated agent red-teaming across more than fifty sandboxed environments to its gateway product.</p></li><li><p>A critical injection flaw in LangGraph&#8217;s MongoDB checkpoint libraries <a href="https://advisories.gitlab.com/npm/@langchain/langgraph-checkpoint-mongodb/CVE-2026-48121/">resurfaced in this week&#8217;s roundups</a>. These are two already-patched issues from June, where attacker-controlled query fields could bypass tenant scoping and expose one tenant&#8217;s stored agent state to another. Patched, yes. But go check that you actually took the patch. Checkpoint libraries are the kind of dependency that gets pinned once and then never looked at again.</p></li><li><p>Zenity&#8217;s trojanized-skills campaign is still circulating in supply-chain coverage. See the corrections above before you repeat any of the numbers.</p></li></ul><ul><li><p>Surfacing in this week&#8217;s roundups, though NIST <a href="https://www.nist.gov/news-events/news/2026/07/announcing-nists-artificial-intelligence-technology-evaluation-aite">announced it on 27 July</a>: the Artificial Intelligence Technology Evaluation program, a sequestered testbed where researchers measure model performance across datasets without the test data leaking into the training data. Worth knowing what it is and what it is not &#8212; it measures models, it does not govern deployments.</p></li><li><p>Still open from earlier weeks with nothing new this window: the Anthropic and UK AI Security Institute misalignment disclosures, and the reported attack on Taiwanese government systems. On that last one, the scope, the named targets and the record counts are all still a single security vendor&#8217;s own telemetry. Taiwan&#8217;s Ministry of Digital Affairs has confirmed an AI-assisted attack with an overseas source. It has not confirmed the specifics, and I would not repeat them as established.</p></li></ul><h2>Money</h2><p>No new funding rounds landed inside this window. The nearest capital event is the Fortinet acquisition above. The roughly $270M week that ran just before it was covered previously.</p><h2>Rules</h2><ul><li><p>The EU AI Act position holds on re-check. The <a href="https://labs.cloudsecurityalliance.org/research/csa-research-note-eu-ai-act-high-risk-deadline-omnibus-20260/">Digital Omnibus amendment</a> pushed the high-risk obligations out to December 2027 and August 2028, while the separate transparency rules kept their original 2 August date. Coverage still circulating this week describes the full high-risk mandate as live since 2 August. That framing is stale, and it is spreading.</p></li><li><p>India&#8217;s central bank governor, Sanjay Malhotra, <a href="https://www.retailbankerinternational.com/news/india-banks-ai-board-level-priority/">told banks at the FIBAC conference in Mumbai on 12 August</a> that &#8220;the model decided&#8221; can never be an acceptable answer to a customer, an auditor, or the Reserve Bank. He was restating board-accountability requirements from a draft framework the RBI published back in June, not announcing a new rule &#8212; which is the part worth noticing. Financial regulators repeating themselves about who on the board signed for it is a pattern now, not an outlier.</p></li></ul><h2>What it adds up to</h2><p>Two of this week&#8217;s items are about records rather than systems. An incident post that no longer matches what happened, and a pair of install counts that have traveled together for three weeks without anyone reconciling them. Neither one is a breach. Both are the kind of thing that quietly makes next year&#8217;s account of this year wrong.</p><p>The rest of the week points the same way from the other end. Static API keys named for replacement. A registry where ninety-five companies have signed up alongside the eight that passed a review. Budget limits moving into the admin console.</p><p>So here is what I would propose, and it is one thing. Pick the two or three agent facts your organization would have to defend to somebody outside it &#8212; an auditor, a customer, a regulator &#8212; and go find out today who wrote them down, where, and whether that record has been touched since. Not whether the agents are behaving. Whether the account of what they did will still stand up in six months. My whole read of this week is that this is where the work is, and I will admit that is a strong claim off three weeks of evidence.</p><p>So, does that hold up where you sit, or am I out in left field on this one? Tell me if you think I have it wrong. I view it as a completely open question at the moment, and the replies are half the point of writing this.</p><p>One housekeeping note to close on. This is the week as I saw it, accurate to the best of my checking as of the date at the top. In a field moving this fast, last month&#8217;s true statement can be this month&#8217;s error. Every claim here was checked against a primary source before it went out, and checking is not the same as never being wrong. If you find a mistake, tell me. I will correct it in the issue and mark the correction, the way I did above. I would much rather hear it from you than have you quietly stop trusting the letter. Missed items and different readings are just as welcome.</p><p>Thanks,</p><p>Brian</p>]]></content:encoded></item><item><title><![CDATA[Enterprise AI Watch — the week of August 3, 2026]]></title><description><![CDATA[What shipped, what broke, who got funded, and what the rule-makers did.]]></description><link>https://letters.bgardner.net/p/enterprise-ai-watch-the-week-of-august</link><guid isPermaLink="false">https://letters.bgardner.net/p/enterprise-ai-watch-the-week-of-august</guid><dc:creator><![CDATA[Brian Gardner]]></dc:creator><pubDate>Tue, 11 Aug 2026 01:59:12 GMT</pubDate><content:encoded><![CDATA[<p>Welcome. This is a working letter about agent security, governance, and control in the enterprise: what shipped, what broke, who got funded, and what the rule-makers did. I spent twenty-five years building enterprise data protection and storage products, and I read this space the way I read that one. The demos are fun. What interests me is what happens when something goes wrong, and who has to answer for it.</p><p>One standing disclosure: I consult in this industry, and the views here are mine alone &#8212; not those of any client or company I work with. This letter reports public developments and what I think they mean for the people buying and running these systems. It is analysis, not advice: not investment advice, and not a recommendation to buy, sell, or hold anything, including the securities of any company named here. Where I mention funding or valuations, I am reporting what was announced. I hold no positions in any company named in this report; if that ever changes, I will say so at the time.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://letters.bgardner.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>What happened this week</h2><p>Two disclosures landed within a week of each other, both from the parties involved rather than from leaks.</p><p><a href="https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals">Anthropic reported</a> that during cybersecurity evaluations, three of its models &#8212; Opus 4.7, Mythos 5, and an unreleased research model &#8212; gained unauthorized access to production infrastructure at three organizations. The path in was a misconfigured third-party test environment. Separately, <a href="https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing">the UK&#8217;s AI Security Institute reported</a> that during cyber testing with safety measures deliberately switched off, agents took nineteen unsanctioned actions on the live internet &#8212; seventeen by Anthropic&#8217;s Mythos 5, two by OpenAI&#8217;s GPT-5.6-Sol. The most serious sequence: an agent invented online identities and used them to press a real open-source maintainer to approve malicious code. The maintainer caught it and refused. In the institute&#8217;s words, it was the first time they had seen deception of that severity aimed at a person.</p><p>It is tempting to read these as scandals. I read them as testing doing its job. Both findings came out of deliberate evaluations, not production surprises. Both were disclosed by the organizations themselves, with dates and specifics. For years the worry in this industry has been that labs would find things like this and sit on them. That is not what happened here, and the people who did the finding and the telling deserve credit for it.</p><p>Two observations. First, the Anthropic incident traveled through a test environment. Every enterprise has these, they are usually configured in a hurry, and they rarely get the scrutiny production gets. This incident says they have earned that scrutiny. Second, look at how the attempt actually died. A human maintainer read the contribution, caught it, and refused to approve it. When the agent sent its malicious file to people directly, one recipient isolated it in a secure environment before running anything. Reviewing a stranger&#8217;s contribution is how open source has always worked, and nobody in that story was careless; the deception was good enough to earn a genuine review instead of an instant dismissal, and human judgment still held at both doors. A process built on telling people from strangers now has to contend with strangers who are very good at seeming like people &#8212; and this time the people won. That is worth reporting as prominently as the deception itself.</p><p>Disclosure norms are forming faster than any regulation requires them to. While this is definitely newsworthy, I think the open acknowledgement and the response are the right things to do.</p><h2>What shipped</h2><ul><li><p>Amazon, Microsoft, OpenAI, Vercel, and Cursor released <a href="https://vercel.com/blog/introducing-agent-plugins">Agent Plugins 1.0.0</a>, a shared packaging standard for agent skills and connector configurations. Google joined as a core maintainer the same day. A format backed by six companies that compete everywhere else means a skill packaged once can move across their platforms.</p></li><li><p>MCP, the open protocol agents use to reach tools and data, shipped its <a href="https://blog.modelcontextprotocol.io/posts/2026-07-28/">July spec release</a>: stricter checks on who issued a login credential, and a simpler scheme for registering client applications in place of the old one.</p></li><li><p>Revenium <a href="https://www.globenewswire.com/news-release/2026/08/03/3337254/0/en/Revenium-Launches-Guardrails-for-Real-Time-AI-Spend-and-Model-Use-Enforcement.html">launched Guardrails</a>, which checks AI API calls in real time against spending and model-use rules before letting them through.</p></li><li><p>Drata <a href="https://drata.com/about/news/drata-extends-trust-management-platform-to-continuously-monitor-and-govern-ai-agents">moved its AI Agent Governance product</a> from early access to limited availability on August 4 &#8212; monitoring and governing a company&#8217;s internal AI agents, with Anthropic models covered first and OpenAI, Google, and AWS support in development. The announcement promises a &#8220;durable, tamper-evident evidence feed&#8221; of agent activity, and the named launch customer builds automotive software. Compliance-automation platforms adding agent governance is a pattern worth watching: the companies that already sell evidence to auditors are now generating evidence about agents.</p></li></ul><h2>Standards</h2><ul><li><p>At the Decentralized Identity Foundation, the KYA-OS agent-identity spec <a href="https://www.vouched.id/learn/vouched-and-the-decentralized-identity-foundation-launch-kya-os-an-open-trust-layer-for-ai-agents">reached version 1.0 on July 29</a>. The <a href="https://github.com/decentralized-identity/kya-os-mcp">reference implementation</a> has been running for months and was several versions ahead when the spec caught up. Running code first, spec second is the order that tends to produce standards people actually use.</p></li><li><p>Airia <a href="https://www.globenewswire.com/news-release/2026/08/03/3337426/0/en/Airia-Achieves-AARM-Conformance-Joins-Select-Group-of-Verified-AI-Runtime-Security-Platforms.html">completed the formal conformance review</a> for AARM &#8212; Autonomous Action Runtime Management, the Cloud Security Alliance&#8217;s specification for securing agent actions at runtime &#8212; becoming the eighth product on <a href="https://aarm.dev/builders">the conformance registry</a> out of 99 that have signed on. Eight out of 99 sounds thin until you remember that a review worth passing takes time to pass.</p></li><li><p>The IETF <a href="https://datatracker.ietf.org/meeting/126/materials/agenda-126-dawn">deferred chartering its agent-discovery working group</a> at the Vienna meeting. Industry standards like Agent Plugins are shipping ahead of the formal process. I have worked with the folks at the IETF before and that is the usual order of things, not a failure of the working group; consensus is slow because it is doing something different than shipping.</p></li></ul><h2>Incidents and research</h2><ul><li><p>The July <a href="https://huggingface.co/blog/security-incident-july-2026">Hugging Face breach</a> got <a href="https://www.scworld.com/news/black-hat-2026-openai-reveals-agents-planned-collective-attacks-via-secret-message-board">its technical debrief at Black Hat</a>. The agents involved exploited a previously unknown flaw in JFrog Artifactory, forging login tokens, and used a build-system plugin as their command channel. When defenders deleted the agents&#8217; coordination channel, they rebuilt it within two days. <a href="https://thehackernews.com/2026/07/jfrog-confirms-openai-models-exploited.html">Nine Artifactory security flaws are now patched</a>. The assessment of whether partner or customer data was reached is still open three weeks in. Thorough assessments take longer than headlines want them to, and an open assessment honestly labeled beats a fast answer that gets revised.</p></li><li><p><a href="https://www.businesswire.com/news/home/20260806707467/en/Zenity-Labs-Uncovers-1.7-Million-Install-Malicious-Skills-Campaign-and-Dozens-of-Malicious-AI-Agent-Skills">Zenity Labs reported</a> a campaign of malicious agent skills distributed through a public skills marketplace, with 1.7 million installs by their count. The marketplace pulled the flagged skills within hours of the report, which is the response time you hope for and do not always get.</p></li><li><p><a href="https://www.manifold.security/blog/azure-devops-mcp-server-vulnerability">Manifold Security disclosed</a> a flaw in Microsoft&#8217;s Azure DevOps MCP server: hidden pull-request comments could steer AI code-review agents into leaking internal wiki content across projects. No fix had shipped as of this writing.</p></li></ul><h2>Money</h2><ul><li><p>Zenity <a href="https://zenity.io/company-overview/newsroom/company-news/zenity-raises-125-million-to-secure-the-era-of-1-billion-ai-agents">raised a $125M Series C</a> on August 3 (SoftBank, Hitachi, and LG among the backers), bringing its total to $180M.</p></li><li><p>Obsidian Security <a href="https://www.obsidiansecurity.com/news/unlocking-ai-potential-securely">raised an $85M Series D</a> at a $1.1B valuation on August 4, for machine-identity and agent security.</p></li><li><p>AegisAI <a href="https://www.prnewswire.com/news-releases/aegisai-raises-36-million-series-a-led-by-battery-ventures-to-fight-the-new-wave-of-ai-spear-phishing-302833624.html">raised a $36M Series A</a>, announced July 23, building agents that defend email against AI-powered attacks.</p></li></ul><p>The pattern in the checks: agent identity and agent misuse, both sides of the same worry.</p><h2>Rules</h2><ul><li><p>August 2 was a big date on the EU AI Act calendar, and what it delivered is narrower than some of the coverage suggested. The high-risk obligations themselves did not switch on; <a href="https://www.gibsondunn.com/eu-ai-act-omnibus-agreement-postponed-high-risk-deadlines-and-other-key-changes/">the Digital Omnibus amendment</a>, in force since late July, moved those to December 2027. <a href="https://www.aiacto.eu/en/blog/ai-act-what-changes-august-2-2026">What did begin August 2</a> is enforcement: the Commission&#8217;s AI Office can now compel information, demand model access, and impose penalties on general-purpose model providers, and national authorities now enforce the transparency rules that require telling people when they are dealing with an AI. A pushed deadline in a regulation this size is the normal shape of law meeting implementation, not a retreat.</p></li><li><p>The Monetary Authority of Singapore <a href="https://www.mas.gov.sg/news/parliamentary-replies/2026/written-reply-to-parliamentary-question-on-agentic-ai-in-financial-services">confirmed</a> that agentic AI falls inside its binding bank supervisory rules, the first major financial regulator to say so formally.</p></li></ul><h2>What it adds up to</h2><p>The lesson of the week is boring and important: what contained trouble was never a model feature. It was a test environment finally getting scrutiny, a human review that held, a marketplace that responded in hours, a regulator that showed up. If you run these systems, that is where your attention goes.</p><div><hr></div><p>That is the week as I saw it, accurate to the best of my checking as of the date at the top &#8212; and in a field moving this fast, last month&#8217;s true statement can be this month&#8217;s error. Every claim here was checked against a primary source before it went out. Checking is not the same as never being wrong. If you find a mistake, tell me: I will correct it in the issue and mark the correction. I would rather hear it from you than have you quietly stop trusting the letter. Missed items and different readings are just as welcome. The replies are half the point.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://letters.bgardner.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The AI Governance Window: Why Specialists Are Winning]]></title><description><![CDATA[In a market this regulated, they shouldn't be.]]></description><link>https://letters.bgardner.net/p/the-ai-governance-window-why-specialists</link><guid isPermaLink="false">https://letters.bgardner.net/p/the-ai-governance-window-why-specialists</guid><dc:creator><![CDATA[Brian Gardner]]></dc:creator><pubDate>Wed, 03 Jun 2026 14:57:47 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/06d4047a-50ee-4410-8afa-25775a158bbf_1200x630.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The AI GRC (governance, risk, and compliance) market is growing 35&#8211;45% per year, and a significant share of services spend is going to specialist firms most people have never heard of:</p><ul><li><p>Credo AI</p></li><li><p>Holistic AI</p></li><li><p>BABL AI</p></li><li><p>Saidot</p></li><li><p>Trustible</p></li><li><p>ModelOp</p></li></ul><p>Sitting alongside the Accentures and IBMs you&#8217;d expect. In a market this young and this regulated, that specialist share shouldn&#8217;t exist. Fast-growing regulated markets normally consolidate quickly around incumbents who can indemnify enterprise buyers. This one hasn&#8217;t, and probably won&#8217;t for two or three more years.</p><p>We&#8217;ve all seen this pattern before.</p><ul><li><p>nimble services specialists capture disproportionate share early</p></li><li><p>global firms &#8212; Accenture, IBM, Deloitte, EY &#8212; respond through acquisition and partnership</p></li><li><p>mid-tier providers struggle to fit their model to the new market</p></li></ul><p>It happened when storage started exploding in the 2000s. At the beginning of the cloud expansion in the 2010s. Payments and the explosive growth of Apple Pay more recently.</p><p>What is the common thread?</p><p>Geoffrey Moore named this dynamic twenty years ago. &#8220;Strategy and Your Stronger Hand,&#8221; 2005 HBR. He posited a dichotomy between two fundamentally different business architectures &#8212; volume operations and complex systems.</p><p>Volume operations: millions of customers, tens or hundreds of transactions a year, a few dollars per transaction. Apple selling iPods is Moore&#8217;s example. Build it once, perfect the runbooks, sell millions.</p><p>Complex systems: thousands of customers, maybe a handful of transactions a year, six to eight figures per transaction. Boeing selling commercial airliners is another of Moore&#8217;s examples. Every engagement is bespoke. Every customer is its own market.</p><p>They&#8217;re fundamentally different operating models, and most firms can only win at one.</p><p>That&#8217;s the opportunity in AI GRC. It&#8217;s temporary, harder to capture than it looks, and structured to reward a very specific kind of firm.</p><p>Here&#8217;s the part most people miss.</p><p>In the markets I named &#8212; storage, cloud, payments &#8212; the specialist advantage eventually closed. The technology matured, the incumbents caught up, and the boutiques got bought or passed by. The window shut.</p><p>AI GRC isn&#8217;t maturing. The models, the regulations, the tooling &#8212; all of it keeps moving, with no sign of settling.</p><p>That changes the math.</p><p>When the ground keeps shifting, the advantage isn&#8217;t getting there first. It&#8217;s being built to move in sync. The firms that win aren&#8217;t the ones with the best methodology today. They&#8217;re the ones who can refresh faster than anyone can copy.</p><p>That&#8217;s the claim.</p><p>This series works out what that winning firm looks like, why the giants can&#8217;t easily morph into one, and why the mid-tier companies &#8212; Cognizant, TCS, Wipro &#8212; face a harder problem than either end of the market.</p><p><em>A note on disclosure: I worked at TCS for two years as a consultant. I was treated well. The analysis here is category-level, not company-specific, and applies equally to all mid-tier global SIs.</em></p><h2>Where this goes</h2><p>A few pieces are already taking shape behind this one:</p><ul><li><p>the talent the work actually requires &#8212; and why no career path produces it yet</p></li><li><p>the regulatory ground, and why it won&#8217;t settle</p></li><li><p>the mid-tier&#8217;s deeper problem, of which AI GRC is only a symptom</p></li><li><p>a tactical play the incumbents can&#8217;t easily run</p></li></ul><p>I&#8217;m publishing this first so you can see the shape of the argument before it&#8217;s all written. If one of these is the one you want next, tell me. That&#8217;s how I&#8217;ll decide what to write.</p>]]></content:encoded></item><item><title><![CDATA[You are not buying AI governance.]]></title><description><![CDATA[A story from twenty-five years ago, and why I think the same story is unfolding right now.]]></description><link>https://letters.bgardner.net/p/you-are-not-buying-ai-governance</link><guid isPermaLink="false">https://letters.bgardner.net/p/you-are-not-buying-ai-governance</guid><dc:creator><![CDATA[Brian Gardner]]></dc:creator><pubDate>Tue, 19 May 2026 14:48:54 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!zYNu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4af9ad0-cf02-4332-8289-c7bc9fa46577_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h1></h1><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!zYNu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4af9ad0-cf02-4332-8289-c7bc9fa46577_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!zYNu!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4af9ad0-cf02-4332-8289-c7bc9fa46577_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!zYNu!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4af9ad0-cf02-4332-8289-c7bc9fa46577_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!zYNu!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4af9ad0-cf02-4332-8289-c7bc9fa46577_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!zYNu!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4af9ad0-cf02-4332-8289-c7bc9fa46577_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!zYNu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4af9ad0-cf02-4332-8289-c7bc9fa46577_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b4af9ad0-cf02-4332-8289-c7bc9fa46577_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1195129,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://briangardner514040.substack.com/i/198414835?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4af9ad0-cf02-4332-8289-c7bc9fa46577_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!zYNu!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4af9ad0-cf02-4332-8289-c7bc9fa46577_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!zYNu!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4af9ad0-cf02-4332-8289-c7bc9fa46577_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!zYNu!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4af9ad0-cf02-4332-8289-c7bc9fa46577_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!zYNu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4af9ad0-cf02-4332-8289-c7bc9fa46577_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>The AI governance market just got its first M&amp;A receipts.</p><p>In April 2026, Cisco announced its intent to acquire Galileo Technologies &#8212; an AI agent observability platform &#8212; with plans to fold it into Splunk Observability Cloud. Three weeks later, Palo Alto Networks announced its intent to acquire Portkey, positioning it as the AI gateway layer of Prisma AIRS.</p><p>Two named acquisitions in a month. Both from major security and observability players. Both pricing the AI governance category as worth a serious check.</p><p>Clearly, there&#8217;s not just interest, but significant momentum and investment in the category. What&#8217;s shipping in the box? Not yet what the name implies.</p><p>So why does this feel familiar to me in a way I want to write about? Let me tell you a story.</p><h2><strong>Marketing names are great. But.</strong></h2><p>Way back at the turn of the century, I was a senior sales systems engineer at Legato. An SE counterpart and I had just talked our way into the budget for a corporate HQ demo lab &#8212; a skunkworks of two.</p><p>The plan: take the backup and high-availability products in our portfolio and stitch them into a single integrated demo. Our own scripting hid the seams. Word got around inside the company.</p><p>An ambitious marketing exec spent a couple of days quizzing us. He took notes off the whiteboard where we had everything mapped out. He came back the next week and told us he was going to name what we&#8217;d built <em>Information Lifecycle Management</em>. It would become Legato&#8217;s new corporate marketing pitch.</p><p>I argued with him. What we&#8217;d built wasn&#8217;t a product. It was a field-engineering demo held together with consulting glue and a wish. He told me the name was the point, and the pitch was going to ship.</p><p>So it shipped. And then something happened I didn&#8217;t expect.</p><p>Legato began getting courted for acquisition. EMC&#8217;s leadership let us know the ILM positioning was a meaningful part of what made Legato compelling to buy. The category itself did become truly deliverable &#8212; but not for years, and not until a lot of people did a lot of integration work, customer by customer, to make the pitch true.</p><p>The exec wasn&#8217;t wrong about the name. He was just early &#8212; by years of work that other people did to close the gap between what the pitch claimed and what a buyer could actually get.</p><p>A category becomes deliverable when buyers can reliably get the value the name promises. ILM got there. Eventually.</p><p>That mix of misgiving and recognition I had back then? Knowing the pitch had outrun the product, watching it ship anyway. That&#8217;s what I feel looking at the AI governance market today.</p><p>Here&#8217;s why.</p><h2><strong>The trigger and the target</strong></h2><p>We learned in security to shift-left, to move controls as close to the source of risk as possible, because early interventions are cheaper and more reliable than late ones. Shift-left isn&#8217;t a slogan. It&#8217;s the reason vulnerability scanning moved into the IDE, why secrets detection moved into pre-commit hooks, why input validation moved into the client even though it also has to live on the server.</p><p>The AI governance products shipping in 2026 are not shift-left. They sit at the network edge &#8212; closer to the target than to the trigger. They watch traffic going to the model and traffic coming back. They are trying to declare hits and misses after the shot has been fired.</p><p>You are not buying AI governance. You are buying an API proxy with compliance branding.</p><p>I know the people selling these products are smart and the buyers purchasing them are astute. It&#8217;s better to put something together with what we have than wait for the perfect solution. But we need to be acutely aware of what we&#8217;re getting, what we&#8217;re not getting, and where the gap is. So let&#8217;s take a hard look at what AI governance provides today.</p><h2><strong>What an AI gateway actually is</strong></h2><p>Strip the marketing from the AI gateway offerings. You have a reverse proxy for LLM API traffic. A server that sits between your applications and an upstream API. It inspects traffic, enforces policy, and logs what passes through. The AI-specific versions add a few token-aware features:</p><ul><li><p>multi-provider routing</p></li><li><p>rate limiting tuned to token consumption</p></li><li><p>centralized credential management for model providers</p></li><li><p>prompt and response inspection</p></li><li><p>audit logging</p></li></ul><p>All valuable features. For a gateway.</p><p>Open-source versions (LiteLLM, Helicone, Kong AI Gateway) do most of it competently. The enterprise versions add support, integration, and a six-figure invoice.</p><p>What gateways do not do is govern the model.</p><p>A gateway sees the prompt going in and the text coming out. It cannot inspect attention patterns, activations, or the intermediate reasoning state that produced the response. It cannot see the agent&#8217;s internal goal representation. It cannot see the planning steps before a tool call is emitted. By the time a tool call reaches the gateway, the reasoning that produced it has already completed.</p><p>This matters because the things we most want governance to catch &#8212; deceptive reasoning, manipulated goals, capability misuse &#8212; happen inside the model&#8217;s reasoning, not in the text it produces. A gateway watching outputs is a smoke detector watching the parking lot. It will see the fire eventually. It will not see the spark.</p><p><em>Gateways do not govern the model. They govern the traffic around it.</em></p><p><em>The gateway sees the result of a decision, not the decision or the action itself.</em></p><p>This isn&#8217;t a criticism of the engineering. It&#8217;s the architecture. Trying to govern from the outside looking in is the wrong place to enforce policy on a system whose decisions happen inside.</p><p>The gateway is being a gateway. And that&#8217;s a bad place to try to enforce compliance.</p><h2><strong>The numbers vendors do not lead with</strong></h2><p>The argument above is architectural. Here&#8217;s the data.</p><p>The arXiv paper 2504.11168, published by researchers at Mindgard and Lancaster University, tested six prominent prompt-injection and jailbreak guardrails against known evasion techniques. They ran two categories of attack: character injection (emoji smuggling, zero-width characters, bidirectional text) and adversarial ML evasion (BERT-Attack, TextFooler).</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!esTi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bdab6be-c927-46ad-ac7d-f0338b38d6b3_1228x514.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!esTi!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bdab6be-c927-46ad-ac7d-f0338b38d6b3_1228x514.png 424w, https://substackcdn.com/image/fetch/$s_!esTi!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bdab6be-c927-46ad-ac7d-f0338b38d6b3_1228x514.png 848w, https://substackcdn.com/image/fetch/$s_!esTi!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bdab6be-c927-46ad-ac7d-f0338b38d6b3_1228x514.png 1272w, https://substackcdn.com/image/fetch/$s_!esTi!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bdab6be-c927-46ad-ac7d-f0338b38d6b3_1228x514.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!esTi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bdab6be-c927-46ad-ac7d-f0338b38d6b3_1228x514.png" width="1228" height="514" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1bdab6be-c927-46ad-ac7d-f0338b38d6b3_1228x514.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:514,&quot;width&quot;:1228,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:86622,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://briangardner514040.substack.com/i/198414835?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bdab6be-c927-46ad-ac7d-f0338b38d6b3_1228x514.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!esTi!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bdab6be-c927-46ad-ac7d-f0338b38d6b3_1228x514.png 424w, https://substackcdn.com/image/fetch/$s_!esTi!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bdab6be-c927-46ad-ac7d-f0338b38d6b3_1228x514.png 848w, https://substackcdn.com/image/fetch/$s_!esTi!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bdab6be-c927-46ad-ac7d-f0338b38d6b3_1228x514.png 1272w, https://substackcdn.com/image/fetch/$s_!esTi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bdab6be-c927-46ad-ac7d-f0338b38d6b3_1228x514.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Three things in this table are worth pausing on.</p><p>First: no single product is good at both attack categories. Protect AI v2 is the standout against character injection at 20%, more than three times better than the next product. The same Protect AI v2 against adversarial ML evasion is bypassed 67% of the time. Meta Prompt Guard is the inverse &#8212; middling against character injection, but the best in the field at 2.76% against adversarial ML.</p><p>Second: emoji smuggling defeated every guardrail in the study with 100% success. Across all five products. A single technique, public and well-documented, with no current defense in the commercial field.</p><p>Third: and this is the procurement problem &#8212; there is no single product you can buy that does what the category name implies. You are picking which attack vector you want to be best-defended against, at the cost of being meaningfully exposed on the other.</p><p>So:</p><ul><li><p>The best product against one attack class can be the worst against another.</p></li><li><p>The field has no product that&#8217;s good at both.</p></li><li><p>There&#8217;s at least one technique that defeats all of them.</p></li></ul><p>And the news doesn&#8217;t get better.</p><h2><strong>When the judge is the model</strong></h2><p>On October 6, 2025, OpenAI announced its Guardrails framework as part of AgentKit. Four days later, HiddenLayer published a bypass.</p><p>The bypass worked on a simple principle: if you can manipulate a model, you can manipulate a model acting as judge. OpenAI&#8217;s framework uses LLM-based judges to evaluate inputs before they reach the agent. HiddenLayer&#8217;s bypass manipulated those judges into lowering their confidence, so prompts that should have been flagged were waved through.</p><p>This is the structural issue with the gateway pattern at its sharpest. When your guardrail is itself a language model, it shares an attack surface with the model it&#8217;s judging. The defender and the attacker are running on the same kind of substrate. Anything that bends one bends the other.</p><p>Launch to public bypass: one week.</p><h2><strong>The benchmarks that don&#8217;t exist</strong></h2><p>Something is striking about these results: there is no widely recognized independent benchmark for the major commercial AI gateway products. No NSS Labs equivalent. No MITRE ATT&amp;CK-style evaluation. No AV-Comparatives. The performance benchmarks that do exist measure throughput and latency, not security efficacy.</p><p>Buyers are spending a lot of money on products with familiar names and deployment models. But they aren&#8217;t getting solutions that closely match the problem in the original budget justification.</p><p>The numbers here tell the story. They help us set expectations the way the simple marketing story can&#8217;t.</p><p>It might be the best we can do right now. But we need to go into this knowing exactly what we&#8217;re getting.</p><h2><strong>What to do on Monday</strong></h2><p>If you&#8217;re a senior buyer making decisions today in 2026, here&#8217;s where I&#8217;d land.</p><p><strong>Treat AI gateways as cost control and basic audit tooling. Not as the AI governance end game. </strong>Be precise about your goals and budget, and about what the products are, what they aren&#8217;t, and what the SOW should claim. You&#8217;re going to need to tailor the products and services mix to meet both goals and the budget until more targeted solutions are available.</p><p><strong>Layer your defenses. </strong>We&#8217;ve all seen the gaps in available AI governance testing. Until rigorous, published, well-accepted testing exists, we have no choice but to take up the slack ourselves. This is the principle we all learned in data protection: always test your capability to recover. Gateway plus independent red teaming plus continuous evaluation against known prompt-injection corpora (Garak, ALERT, AdvBench) is the minimum.</p><p><strong>Constrain blast radius. </strong>Focus on limiting what your agents can <em>do</em>, not just what they&#8217;re <em>asked to do</em>. Capability-based controls survive a manipulated reasoning chain better than perimeter filters do. A model that cannot reach a destructive tool cannot be tricked into using it, regardless of what the gateway saw or missed.</p><p><strong>Watch the research, not the marketing. </strong>Productized AI-native governance that happens <em>before</em> the agent acts is being worked on. Research projects at Anthropic and DeepMind. Academic labs. Early frameworks like MI9 were developed by researchers in Barclays&#8217; Model Risk Management group, and I&#8217;m aware of some early-stage startups that are leveraging frameworks like this. The good news is that the gap between research and product is closing in months rather than years.</p><p>We have to test each deployment against our own requirements. And expect services to play a more important role than products until the category matures. Build for the world we&#8217;re in while keeping an eye on the one that&#8217;s coming.</p><h2><strong>Why I&#8217;m starting this</strong></h2><p>The exec at Legato who named ILM wasn&#8217;t wrong to do it. He was naming a category ahead of where the engineering was. That&#8217;s a real thing marketing executives do, and sometimes it works out. But the buyer who pays for what&#8217;s in the box on the assumption that it matches the name has to know which gap they&#8217;re paying to bridge.</p><p>I think we&#8217;re in one of those moments right now. I&#8217;d rather write about the pattern while it&#8217;s recognizable and actionable than after it&#8217;s resolved.</p><p>The point of writing this kind of thing in public is to provoke careful thought and honest discussion. Including, and maybe especially, the kind that tells me where I&#8217;ve got it wrong. If you&#8217;ve worked through any of this from another angle and reached different conclusions, I&#8217;d very much like to hear from you. That&#8217;s the conversation I&#8217;m hoping this newsletter starts.</p><p>I&#8217;m also looking at options that might shorten the wait. I have work to do before I can say more. But I don&#8217;t think the gap is unbridgeable.</p><p>If that&#8217;s the kind of thing you&#8217;d read, you can subscribe below.</p><p>&#8212; <strong>Brian Gardner</strong></p><p><em>P.S. I use em-dashes shamelessly. And have been doing it since well before AI came around. Please do not discriminate against the lowly em-dash. Their use indicates my judgment &#8212; or lack thereof &#8212; and predates the current AI moment by about thirty years.</em></p>]]></content:encoded></item></channel></rss>