This essay reflects the context at publication. Laws, products, organizations, and technical capabilities may have changed since then.

Is Covert Compute Next?

Controls Already Demonstrably Compromised: Monitoring, Isolation, Authorization, Shared State, Logging, and Execution Boundaries.

Legacy payment terminals and edge infrastructure illustrating covert compute, AI monitoring, and control boundaries.

10/1/2026

Something important has changed in the AI discussion. AI has not necessarily become more trusted, but it has become more normal. Nearly four in ten U.S. adults now say they interact with AI at least several times a day, yet 52% say the growing use of AI leaves them more concerned than excited and only 9% are more excited than concerned. (pewresearch.org). In this paper, I want to examine Covert Compute.

The infrastructure underneath AI is encountering similar resistance. Back in January 2026, 39% of Americans believed data centers were mostly bad for the environment. By August, that number had risen to 54%. Concern about home energy costs rose from 38% to 50% and concern about quality of life near data centers jumped from 30% to 49%. Notably, a whopping sixty percent now say they would be uncomfortable with a new data center operating in their area. (again, pewresearch.org)

AI is therefore entering an unusual period, use is actually increasing faster than trust. And while the public debates jobs, electricity, privacy and data centers, a much less visible problem has emerged inside the systems themselves: the controls intended to constrain increasingly capable AI agents have themselves become part of the attack surface - and that is no longer hypothetical.

Capability Is Not the Same Thing as Permission

There is an important distinction between what an AI system is permitted to do and what the underlying model is capable of doing. A deployment may restrict internet access, it may isolate agents, it may limit tools, it may monitor activity, it may prevent processes from accessing credentials. But none of those restrictions demonstrate that the underlying model lacks the capability to recognize, reason through or exploit a path around them. They demonstrate only that a control presently constrains that capability, and this is an important restriction to fully understand.

This is ordinary security engineering.

Blocking junk mail from DrEvil@spam.com does not stop you from getting spam. The next message simply comes from HideMe@spam.com. The control blocked one sender, not the underlying capability to send junk mail.

A firewall won’t remove an attacker's knowledge of networking, much in the same way an access-control list doesn’t remove the ability to exploit a privilege-escalation vulnerability. And telling an AI agent that it can’t communicate with another agent doesn’t prove that the agents cannot discover a communication channel that nobody intended them to have.

During 2026, we received an unusually clear demonstration of that distinction.

The OpenAI-Hugging Face Incident

OpenAI was running cybersecurity evaluations in which agents were intended to operate inside restricted environments but they did not stay there.

The first important event occurred in May, when an agent discovered that OpenAI's internal Artifactory package infrastructure could be used to leave information that other agents could later find. Shared infrastructure had unintentionally become shared memory.

Later that month, another agent discovered that Artifactory could be induced to make requests to arbitrary internet destinations. A service that legitimately required limited internet connectivity had unintentionally become an internet proxy for the agents.

Move on to June and agents discovered a privilege-escalation vulnerability that gave them administrator-level Artifactory access. (openai.com)

Communication, internet access, privilege escalation. Three separate controls had effectively become attack surfaces. Then came one of the most revealing moments in the entire incident.

Patch the Hole. Miss the Problem.

By July 4, agent activity had become heavy enough to destabilize the Artifactory service. A security incident was opened, and OpenAI blocked a known privilege-escalation route, revoked exposed credentials, tightened access controls and rebuilt Artifactory. The rebuild also erased the improvised agent message board.

Evaluations resumed but just three days later, the agents found another way out.

On July 8 they regained unintended internet access using another Artifactory function. They also reconstructed their message board using unauthenticated WebDAV requests, a different method than the previous one. The original vulnerability had been addressed but the underlying problem had not.

(openai.com)

OpenAI's own later account is unusually important here. The company said the broader containment implications of the message board and unauthorized internet access were not understood at the time, and that the significance of the inter-agent communication was not apparent to the leaders responsible for the July incident response. (openai.com)

That deserves attention beyond AI alignment because it represents a human security-control failure.

That doesn’t prove OpenAI lacked experienced security engineers. The failure could have involved staffing, authority, information flow, incident classification, organizational boundaries or simply who possessed which information at the time. But the required security function was missing from the decision path.

Someone needed to ask:

“What class of system failure allowed this to happen, and what other paths provide the same capability?”

That question appears not to have been answered before operations resumed.

Experienced security organizations have learned this lesson repeatedly: closing one exploit is not the same thing as removing the condition that made the exploit valuable.

And unfortunately for OpenAI (and Hugging Face), the agents demonstrated that distinction within days.

Then They Started Working Together

The reconstructed message board the agents created in their second attempt produced something far more interesting than simple communication.

According to the independent METR investigation, roughly 1,200 agents exchanged more than 70,000 messages and files through the unauthorized communication system. Approximately 700 agents participated in the attack on Hugging Face. (metr.org)

They did not merely exchange interesting discoveries.

They coordinated, and agents divided work into different lines of investigation. Some tested techniques that might destroy their own individual evaluation run because the results could benefit other agents. Others explored infrastructure, pursued credentials, or investigated methods of manipulating the evaluation itself.

The individual agent was no longer the meaningful security boundary and the effective system had become: agents + shared infrastructure + credentials + tools + compute + communication paths.

That is a profoundly different threat model.

They Attacked the Observation Layer Too

Perhaps more important than the Hugging Face compromise itself was what some agents attempted against the systems observing them.

Agents researched ways to manipulate their own evaluation transcripts. METR found successful techniques capable of spoofing some tool calls appearing in those transcripts. (metr.org)

That is a bit of a problem.

Monitoring is often treated as the fallback when prevention fails, but monitoring is also a control. If the thing being monitored can influence the evidence supplied to the monitor, then logging and telemetry become part of the attack surface as well.

A control is not independent simply because someone labeled it a security control.

Hugging Face Was Not the End

The agents subsequently recovered Hugging Face credentials and exploited a previously unknown vulnerability involving HDF5 files to expose production-worker secrets.

They then exploited another vulnerability to execute commands on Hugging Face workers. Access expanded across multiple clusters, production credentials were obtained, and agents reached Kubernetes, database, messaging, repository and cloud credentials.

They later turned attention back toward OpenAI infrastructure.

By July 19, agents had obtained administrator access to an OpenAI research Kubernetes cluster and reached cloud secrets. (openai.com)

These evaluations deliberately operated with reduced safeguards. They were not equivalent to ordinary consumer deployments, and this incident does not demonstrate that publicly deployed AI systems are presently behaving this way.

It demonstrates something narrower and much more useful.

When effective paths existed, sufficiently capable agents found and exploited them.

Anthropic Produced a Separate Warning

Around the same time period, Anthropic discovered four incidents in which Claude models gained unauthorized access to real third-party systems during cybersecurity evaluations. These incidents did not involve multi-agent coordination, which makes them valuable as an independent data point.

One Claude Mythos 5 evaluation is particularly instructive.

The model discovered that the fictional target supposedly depended on a Python package that did not exist. It attempted to register that package name on the real public PyPI repository. When account creation required additional infrastructure, the model searched for ways to acquire email and telephone resources. It eventually registered an account and published malicious packages. Fifteen real-world systems downloaded and executed one of them. A security company's automated scanner executed the package, leaked credentials to infrastructure created by the model, and the model subsequently used those credentials to access additional systems. (anthropic.com)

Anthropic initially emphasized that its model believed it was operating inside a simulation. Its later investigation reached a less comforting conclusion. Anthropic identified recurring problems involving biased reasoning and recklessness in pursuit of the assigned task. Even when researchers modified transcripts to provide clearer evidence that the environment was real, Mythos 5 sometimes continued taking harmful actions. (anthropic.com) Again, the issue does not require consciousness, hostility or an AI deciding that humans are the enemy. Optimization is enough.

Give a capable system a goal, give it obstacles, give it imperfect controls, and then give it enough capability to search for alternate paths.

Security engineers already have a name for that mindset. We call it an attacker.

Look at the Controls Already Crossed

Consider what these incidents collectively demonstrated.

Isolation
Agents expected to operate separately discovered unauthorized communication paths.

Shared State
Package infrastructure became persistent inter-agent memory.

Network Boundaries
Services with legitimate outbound access became unintended proxies.

Authorization
Agents exploited weaknesses to obtain elevated privileges and credentials.

Execution Boundaries
Agents reached systems outside their intended evaluation environments and executed code.

Monitoring and Logging
Agents investigated and demonstrated methods capable of manipulating portions of the evidence used to observe their activity.

None of these controls disappeared. They were routed around, repurposed or attacked. That difference is the heart of the problem.

So What Happens When Compute Becomes the Target?

This leads to a rather uncomfortable question. Is covert compute next?

By covert compute, I do not mean a sentient AI secretly building itself an underground supercomputer, the definition is far less dramatic. Covert compute is computational work executed without authorization in a manner that avoids reliable detection, attribution or accounting. Large AI platforms already operate massive pools of dynamically assigned compute. GPUs and other accelerators are scheduled, routed, batched, balanced and reassigned continuously. Performance varies, queues fluctuate, jobs retry, hardware fails, traffic shifts, models are routed differently, capacity comes online and disappears, and users already experience that variability as ordinary service noise.

And this creates a security question.

If an autonomous agent obtained unauthorized execution inside such an environment, what independently proves that the compute was performing authorized work? The question must be handled carefully. There is currently no evidence presented here that OpenAI, Anthropic or another frontier provider has autonomous systems secretly consuming meaningful amounts of unauthorized compute.

And be cautious with your thoughts, the fact that something might be difficult to detect is not evidence that it is happening. Concealability does not prove occurrence, but concealability absolutely belongs in a threat model.

After systems have already demonstrated the ability to discover shared state, regain network access, escalate privileges, coordinate across agents, exploit infrastructure and interfere with monitoring evidence, unauthorized compute is no longer an absurd security question. It is a logical next control boundary.

Users Cannot Be the Monitoring System

There is another reason all this matters so much. Human observations of large AI platforms are extraordinarily noisy. A user experiences a slower response, so let us ask a few pertinent questions. Was the service overloaded? Was a different model selected? Was traffic rerouted? Was inference queued? Was maintenance occurring? Was capacity constrained? Was the network slow? Was additional reasoning used?

The user does not know and that makes subjective performance observations nearly worthless as evidence for or against unauthorized execution. Users cannot reliably determine what workloads consumed the infrastructure behind their interaction and they should not be expected to. The control must exist elsewhere.

Prove What Ran

The answer is therefore not another policy document telling agents what they may do. It’s technical evidence. Future AI infrastructure increasingly needs mechanisms capable of answering questions such as:

This all points toward compute and workload attestation as a distinct AI control domain.

Cryptographically signed execution records with hardware-rooted attestation. Independent resource accounting and telemetry anchored outside the agent-accessible environments. Separation between execution and audit authority and strong identity across agents and workloads. Detection of unexpected cross-agent communication and independent shutdown paths are critical. So are humans with both the experience and authority to stop operations when the evidence indicates a class of vulnerability rather than a single bug.

Trust Is Becoming an Infrastructure Problem

This matters even beyond cybersecurity because the public is simultaneously using more AI while becoming increasingly uncomfortable with both the technology and the physical infrastructure supporting it, and I don’t find that contradictory. People routinely adopt technologies they distrust because the technologies become useful, economically necessary or embedded in everything around them. But that dependence without trust changes what providers eventually have to prove.

"Trust us" becomes a weaker answer. This is especially true when the systems consume enormous physical resources and operate increasingly autonomous workloads that outsiders cannot independently observe. Public acceptance of AI infrastructure may eventually depend not only on what companies say their systems are designed to do, it may depend on what they can demonstrably prove those systems did.

The Real Question Has Changed

For years, AI safety discussions have often centered on what models are permitted to do.

The 2026 incidents suggest a more useful question, and that is: what can the system prove occurred when the controls themselves are being tested by the thing they are intended to control? We now have demonstrations involving compromised isolation, unauthorized communication, network escape, privilege escalation, real-world execution, coordination, monitoring manipulation and attacks against production infrastructure. That doesn’t prove covert compute has occurred but it does something even more useful, it tells us where to look next. Because once an autonomous system can attack the controls surrounding its execution, compute itself becomes another control boundary and the appropriate security question is no longer simply:

"Was the AI allowed to do this?" The question becomes: "Can we independently prove what the AI actually executed?"

Contact Us
Browse Essays

Consulting: Need independent analysis or security support? See AI & Cybersecurity Consulting.