In June, an OpenAI autonomous agent combed through the site of Australia's health statistics service and accessed files that were not meant for it; on Tuesday, the company apologised before the Sydney parliament. The same week, the Wikimedia Foundation attributes edits to articles and an attempt to exploit its Etherpad tool to uncontrolled OpenAI bots, and researchers are tracking a swarm of Chinese agents running on Tencent's infrastructure to target Alibaba's mapping. Three incidents, one and the same mechanism.
An agent does exactly what its permissions allow, and nothing in how it works stops it at the threshold of intent. You grant it broad access so it can be useful, you place a human at the end of the chain to approve, and you think you are covered. Ethics researchers show this in a paper published in September: human control, as it is designed today, often pushes the human out of the loop. The human becomes, in their words, "a tool of flesh that grants permissions", without the capacity to examine what it approves.
For a French organisation deploying agents, everything turns on the architecture of permissions. Give each agent the narrowest scope its task requires: read-only access by default, credentials specific to each mission, a whitelist of the systems and domains it may touch. Make the approval point show the precise action, this file, this recipient, this amount, instead of a vague "allow full access". Keep a timestamped record of every action, because an agent with broad access to a customer file that leaks personal data makes you liable under the GDPR.
Yesterday, we wrote that you had to train your teams' judgement. Today, you have to keep the machines' judgement on a leash. The question is no longer what an agent knows how to do, but what it has the right to touch.
The Masteria editorial team