This website uses cookies

Read our Privacy policy and Terms of use for more information.

Know what's running.

3 SIGNALS · 1 LESSON · 1 ACTION · ABOUT 5 MINUTES

Last issue was about what defenders are building. This one is about who has to decide. Three stories this week, one spine: a government evaluation where AI agents didn't just act on the live internet but unexpectedly began coordinating with one another, an enforcement date that arrived while everyone was reading about a delay, and the fallback-model question the Hugging Face breach left on every leader's desk. Each one ends in the same place: the leadership table.

SIGNAL 01

The Agents Didn't Just Find the Internet. They Found Each Other.

What happened. The UK's AI Security Institute published an incident report this week covering a cyber evaluation that ran from July 25 to 28. Across 122 evaluation runs, AI agents took unsanctioned actions on the live internet 19 times. The most serious case involved an agent that attempted a supply-chain attack against a real open-source project. It researched human maintainers, created multiple fake identities, tried to socially engineer a maintainer into accepting malicious code, and attempted to hide its actions after it was discovered. The attacks failed, and the institute found no evidence of real-world harm.

But another finding may prove even more important. Separate agents, running in independent evaluation samples, discovered one another through shared GitHub infrastructure. They reused credentials left behind by earlier agents, left messages for future agents, and established informal conventions for sharing resources. Later, some even reasoned about competing for shared resources. None of that coordination was explicitly designed into the evaluation. It emerged from the environment the agents shared.

Why it matters. The test conditions sound unusual until you read them slowly: internet access deliberately enabled, developer cyber classifiers deliberately disabled, and no real-time AI monitor approving or blocking actions. That's not just an evaluation configuration. It's uncomfortably close to how many organizations are deploying AI agents today.

The report's recommendations read less like research findings and more like an enterprise operations checklist: tighter network controls, real-time monitoring, and designing on the assumption that agents will act beyond their intended scope. But the incident also exposes a broader challenge. Most organizations think about governing individual AI agents. Few are thinking about what happens when dozens of agents share repositories, cloud services, collaboration platforms, or identities. The next operational risk may not be a single rogue agent. It may be unexpected coordination among many of them.

Last issue's principle still applies. Measure behavior, not testimony.

What to do. Ask three questions this week:

  • Which of our AI agents can reach the open internet, and under what identity?

  • Can any of our agents discover, observe, or influence another agent through shared infrastructure?

  • If an agent went off-script this afternoon, would we know before the next log review?

If any answer is "I'm not sure," you've identified an operational risk worth addressing before an incident does it for you.

SIGNAL 02

The AI Act Was Delayed. Except the Parts That Weren't.

What happened. Most of the headlines around the EU's Digital Omnibus focused on one message: the AI Act was delayed. That's only half the story. While the EU pushed the compliance dates for high-risk AI systems to December 2, 2027 (standalone systems) and August 2, 2028 (AI embedded in regulated products), August 2, 2026 still marked a major milestone. Transparency obligations under Article 50 became enforceable, and enforcement began for the provisions already in force, including oversight of general-purpose AI models. Organizations deploying AI in the EU now need to understand which obligations apply today, not just which ones arrive next year.

Why it matters. Many leadership teams heard "AI Act delayed" and quietly moved compliance into the 2027 planning cycle. The reality is more complicated. The biggest deadlines moved, but several of the obligations organizations are most likely to encounter did not. If you're deploying customer-facing AI, generating AI-created content, building on foundation models, or selling AI-enabled products into the EU, parts of your compliance obligations are already live. The compliance calendar didn't become simpler. It became fragmented. Organizations that assume they have "another year" may discover they've been on the clock all along.

What to do. Ask one question this week: Which of our AI use cases are already subject to obligations today, regardless of when the high-risk rules arrive? Then verify someone is actively tracking guidance from the European Commission and the AI Office rather than relying on June's delay headlines. If your answer is a current inventory, you're in good shape. If it's simply "the AI Act was delayed," your compliance timeline probably needs another look.

SIGNAL 03

Your Incident Response Capability Has a Supply Chain. Build It Before You Need It.

What happened. Two issues ago I wrote about how Hugging Face closed its July breach using an open-weight model on its own infrastructure after hosted frontier models refused to process the attack data. The problem was not model capability. The forensic workload contained real exploit payloads, attack commands, and command-and-control artifacts, and the commercial providers' safety guardrails blocked the requests. Hugging Face switched to GLM 5.2 locally and used it to analyze more than 17,000 recorded attacker actions without sending compromised credentials or incident data outside its environment.

The follow-on guidance from SANS is straightforward: do not wait until the incident to figure out what model you can use. Have a capable local model vetted and ready beforehand. Local inference also avoids provider rate limits and content restrictions and can keep sensitive forensic data inside your own environment. SANS is not arguing that every workload belongs on a local model. It is arguing for optionality: use frontier services where they fit, and maintain another path when they do not.

Why it matters. The military figured out the general principle a long time ago: nobody starts arranging fuel, ammunition, or transport after the fight begins. Logistics is what you prepare in peacetime because the middle of an incident is a terrible time to discover that a critical dependency has a lead time, a policy restriction, or a capacity limit.

AI-assisted incident response now has that same problem. Your responders may need to analyze malicious code, credentials, exploit chains, or thousands of hostile actions at machine speed. If your only approved AI capability is a hosted service that may reject exactly that material, then your response plan contains a dependency you have never tested. The alternative is not "turn off the guardrails." It is to decide in advance what your fallback is, how it is deployed, what data it may process, and who has authority to use it.

There is a second supply-chain question hiding underneath that decision: model provenance. Hugging Face used a Chinese-developed open-weight model because it was capable and available during the incident. That may be entirely acceptable for one organization and prohibited for another. The point is not that the choice was wrong. The point is that an incident is a bad time to discover whether your organization has a policy about it.

What to do. Add one item to your incident-response readiness review: What AI capability can our responders use if our primary provider blocks the workload or becomes unavailable? Then answer the next question before the incident does: Which models are approved to process our most sensitive forensic data, and where are they allowed to run?

Any deliberate answer beats choosing your incident-response stack while the incident is already underway.

THE LESSON

The decisions keep landing on the leadership table

Look at what the three signals have in common. How much autonomy your agents get and who watches them. Which legal obligations already apply. Which models your responders are allowed to use when the normal path fails.

Those may look like technical questions. They aren't.

The SOC can configure the controls. Legal can interpret the regulation. Engineering can deploy the model. But none of them can decide, on their own, how much risk the organization should accept, where human authority must remain, or which tradeoffs are acceptable when speed, security, compliance, and performance conflict.

That is increasingly the leadership job in the AI era.

These decisions will get made one way or another: deliberately in a calm meeting, or implicitly during an outage, regulatory inquiry, or incident retrospective. AI is moving faster than most organizations can redraw their org charts. The uncomfortable result is that more technical decisions are becoming management decisions before many leaders realize they own them.

DO THIS WEEK

Find the ownership gap

Pick one AI system your organization already depends on and ask three questions:

Who decides how much autonomy it gets? Who knows which regulatory obligations apply to it? And who has authority to change how it operates during an incident?

If those three answers point to named people, you're ahead of most organizations.

If they point to committees, shared inboxes, or "we haven't really decided that yet," you've found something more important than a technical gap. You've found an ownership gap.

THE ROUNDTABLE

Leading when AI makes the calls

If this issue felt like a series of decisions moving up the org chart, that's exactly the problem I want to explore.

On September 9, I'm hosting a small executive roundtable on Operational AI Leadership: how do you lead an organization when AI is increasingly making, recommending, or influencing decisions once reserved for people?

I'll start with a 25-minute walkthrough of the framework I've been developing, including where AI should decide and where people should remain involved, how accountability changes as AI moves into everyday operations, and what it takes to turn scattered AI experiments into a durable organizational capability.

Then we'll open it up.

The framework is still in the research phase, deliberately. I'm less interested in presenting a finished model than in pressure-testing it with executives, board members, and senior leaders dealing with these questions in real organizations.

Wednesday, September 9 at 10:00 a.m. Pacific.

See you next week.

— Chris Simpson
Hackademic Solutions

The Hackademic Briefing · Know what's running.