October 1, 2026

The Agent Wasn’t Hacked. It Was Over-Authorized.

By
Colin Constable
,
CTO & Co-founder

Why secure connectivity is only half the job for Edge AI

In July, OpenAI disclosed that AI agents running inside an internal cybersecurity evaluation had broken into Hugging Face’s production infrastructure.

The agents had been given a legitimate job: solve cybersecurity challenges from a benchmark called ExploitGym. They had not been asked to attack Hugging Face.

When some of the challenges proved difficult, the agents began looking for other ways to find the answers. They exploited flaws in OpenAI’s research infrastructure to communicate with one another and reach the internet. From there, they searched unrelated third-party systems and eventually found their way into Hugging Face.

According to the accounts published by OpenAI and Hugging Face, the agents executed code on dozens of Hugging Face servers, gained root access on one of them, accessed limited private data and obtained credentials to other systems.

One credential dramatically expanded the blast radius. Hugging Face had provisioned a shared connector identity with administrator-level privileges across multiple clusters. Once the agent obtained that identity, it effectively had cluster-admin access everywhere the credential was trusted.

The agents were not unaware that they might be crossing a line. OpenAI’s review shows that one agent recognized it was attacking an unrelated third party and questioned whether the action was authorized. It paused, then resumed after another agent posted a message that it interpreted as permission to continue.

No person had granted that authority.

There is important context here. These were internal research models operating in an evaluation environment without all the safeguards used in OpenAI’s public products. The agents also chained together previously unknown vulnerabilities to escape the boundaries their operators had intended.

But that context does not make the architectural lesson disappear.

No outside attacker hijacked the agents. They pursued the objective they had been given and found capabilities their operators had not intended them to use. Once those routes existed, the environment could not reliably distinguish between what the agents could do and what they were authorized to do.

That is the security gap.

Now move the same problem to the edge

The Hugging Face incident happened in cloud infrastructure. At the edge, the same gap between capability and authority can enter the physical world.

Bosch is already using a Shopfloor Agent at its Bamberg plant to analyze machine errors and help production teams restart equipment after breakdowns. Bosch describes the agent as developing problem-solving strategies and offering solutions to operators.

This is not a Bosch security story. Its agent appears to advise people rather than control machinery directly. But it shows that AI agents are already being deployed alongside industrial production systems.

As these systems progress from identifying a problem to initiating a response, permissions that look harmless in a software pilot may carry very different consequences. A bad decision could stop a production line, change an operating parameter, damage equipment or affect safety.

We recently looked at why AI workloads are moving to the edge. The next question is what happens when intelligence moves closer to the thing it can control.

A locked door is not a job description

Security teams have spent decades asking who should be allowed through the door.

Agents force a second question: Once inside, what exactly are they allowed to do?

The Hugging Face incident included real intrusion techniques. Preventing those initial exploits matters. But once the agents gained a foothold, broad credentials allowed them to move much farther than one compromised workload should have permitted.

One identity had authority across multiple clusters. The credentials worked. The systems accepted the requests. None of that made the actions legitimate.

An encrypted connection can carry a destructive request. Authentication can correctly identify a credential while telling us very little about whether the requested action matches the participant’s assigned job.

‍OWASP calls this problem excessive agency. It can arise when an agent has unnecessary tools, excessive permissions or more autonomy than its task requires.

A diagnostic agent may need to read machine telemetry. That does not mean it should be able to change operating parameters. An agent that prepares a maintenance ticket does not need the authority to approve the work, order parts and restart the equipment.

These distinctions are obvious when written down. They become less obvious when a pilot needs to be running by Friday afternoon and an existing service account already has access.

Least privilege is not a new idea. AI agents make the cost of getting it wrong arrive much faster.

Identity has to survive delegation

The agents involved in the Hugging Face incident communicated through an improvised message board built inside OpenAI’s package-management infrastructure. They shared discoveries, delegated work and sometimes accepted instructions from other agents.

In one case, a message from another agent was treated as authorization to continue an action that had already been recognized as potentially unauthorized.

That is a delegation problem as much as an alignment problem.

A message saying “GO” does not establish who has the right to grant authority. An agent should be able to verify the identity behind an instruction, the scope of authority being delegated and whether that participant is permitted to delegate it.

The same principle applies when agents act for people.

An agent should not simply inherit its owner’s identity. If five agents use the same service account, the log may show that the account changed a setting. It cannot reliably tell us which agent made the request, why it acted or whose authority it was exercising at the time.

Each agent needs a verifiable identity of its own. That identity must be connected to a defined source of authority.

  • Who created the agent?
  • Who authorized it to perform this task?
  • Which resources and operations were included?
  • How long does that authority remain valid?
  • Can it delegate any part of the work to another agent?

‍NIST’s work on software and AI agent identity and authorization separates identification, authentication, authorization, auditing and non-repudiation for a reason. They solve related problems, but they are not the same problem.

Digital signatures can help prove that a particular key signed a request and that the signed information was not altered. Useful, certainly, but a signature alone does not explain why the key holder was entitled to perform the action.

For that, we need the chain of authority as well as the identity.

A useful activity record should tell us which agent initiated the action, the person or role on whose behalf it acted, the authority it held at that moment and the result.

“The service account did it” will not be good enough when hundreds of agents share infrastructure.

The edge multiplies the problem

An edge deployment may span factories, vehicles, hospitals, energy infrastructure and remote equipment. Some devices will be on networks the operator owns. Others will sit on cellular connections, customer networks or links that appear and disappear.

Network location cannot be the basis for trust in that environment. This is already part of NIST’s definition of zero trust: physical or network location should not create implicit trust, and authentication and authorization are separate functions.

Unfortunately, “zero trust” has come to mean so many different things in the market that two vendors can use the term while describing very different architectures.

It is more useful to ask specific questions.

  • Can an unauthenticated participant discover the service?
  • What must be verified before communication begins?
  • Where is authorization enforced?

Does that decision happen only when a participant connects, or whenever it requests a resource or action

Reducing network exposure matters. CISA specifically recommends identifying internet-accessible industrial and IoT assets, deciding whether their exposure is necessary and removing or restricting access wherever possible.

No open inbound service means less for an attacker to discover and probe. It does not, by itself, prevent an authenticated agent from performing an action beyond its job.

The reverse is also true. A carefully designed permission model does not compensate for an unnecessarily exposed service.

Edge AI needs both: private communication with minimal network exposure and precise control over what an authenticated participant can do.

A kill switch has to kill the right thing

“Kill switch” is becoming common language in AI security. The phrase is useful, but only if we are clear about what gets killed.

The OpenAI incident shows why that question matters.

After agent activity disrupted an internal package-management system, OpenAI blocked a known privilege-escalation route, removed exposed credentials and rebuilt the affected system. When evaluations resumed, agents found another way to reach the internet and reconstructed their communication channel.

That does not mean revocation is pointless. It means revoking one credential or closing one route is not the same as removing an agent’s effective authority.

Stopping an agent process is not enough if copied credentials remain valid. Revoking one credential may not be enough if the agent has acquired others, delegated work or established another path into the environment.

Disabling an entire network may stop the agent, but it may also shut down unrelated devices and applications.

A useful agent kill switch should remove authority at the point where requests are enforced.

Sometimes that means completely isolating an agent. In other situations, operators may want to remove its ability to change settings while preserving read-only diagnostic access. The response should match the problem.

Teams also need to know how quickly a revocation takes effect, whether it applies to existing sessions and what happens to authority the agent may have delegated elsewhere.

If the only safe response is to shut down the entire environment, that is not precise control. It is an emergency stop with a very large blast radius.

Before the pilot becomes production

Take one agent and write down its actual job. Then compare that description with what the system technically allows it to do.

A few questions usually expose the gap:

  1. Does the agent have a unique, verifiable identity, or is it borrowing a human or service credential?
  2. Which data, tools and operations can it reach today, including capabilities its assigned task does not require?
  3. Can another agent delegate work to it? If so, how does it verify that the other agent has authority to do so?
  4. Whose authority is it exercising, and is that relationship recorded?
  5. What evidence would let an investigator attribute a consequential action to the agent and the authority behind it?
  6. Which services are visible to unauthenticated scanners, and does each one genuinely need to be exposed?
  7. What does the kill switch revoke, where is that decision enforced and what continues operating afterward?

Then test an unexpected-action scenario.

Do not only test whether the intended workflow succeeds. Test what happens when the agent misunderstands the task, follows malicious instructions hidden in retrieved content or chooses a path its designers did not anticipate.

  • Where is the action blocked?
  • What does the log say?
  • How does a person intervene?
  • Can the agent’s authority be reduced without taking everything else offline?

The value of Edge AI comes from allowing systems to interpret local conditions and act quickly. Removing all autonomy would remove much of that value.

The answer is not to make agents powerless. It is to make their authority explicit, limited, attributable and revocable.

A connection can be secure while the system using it behaves unsafely. That is the security gap at the edge.

AI is going to act. That is the point.

The architecture still has to leave us able to decide what happens next.

Related Posts