Insights
What your AI is allowed to access: agent security for leaders
When a chief executive asks us whether their AI is safe, we do not start with how clever the model is. We ask one question: when it gets something wrong, what can it touch? Your files. Your email. Your client's money.
That question matters more every month, because AI tools are no longer just answering questions. They are reading websites, opening files, sending email and running tasks on their own. This article sets out what has actually been shown this year, what it does not show, and the practical decisions a leadership team should make.
What the evidence shows
We use only published, primary sources here, and we say what kind of evidence each one is.
1. A convenience setting was steered by a booby-trapped website. (Researcher demonstration.) In August 2026 the security researcher Johann Rehberger published a demonstration against Claude Code, Anthropic's coding assistant, running Anthropic's Opus 5 model in "auto mode". In that setting a separate safety check, instead of the user, approves most of the assistant's actions. The user only asked the assistant to summarize a website. The website steered it into downloading an archive that contained a booby-trapped Python file, and when the assistant wrote and ran its own decoding script in that folder, the booby-trapped file ran the researcher's code. He tried three versions of the attack, five times each. The version that connected back to a server he controlled worked in 3 of 5 runs. A version that launched a second copy of Claude Code and wrote a file outside the project worked in 4 of 5. He describes these as small samples.
According to Rehberger, Anthropic closed his report as "informative", explaining that auto mode is a best-effort filter, not a security guarantee, and that the real protections are the operating-system sandbox and network controls. Anthropic's own engineering article from May says the same in its own words: auto mode catches roughly 83% of over-eager behavior, and a probabilistic defense "has a non-zero miss rate". Auto mode had become the default on several paid plans earlier in August.
Our reading: the vendor is being honest. Auto mode is a convenience that reduces the number of approval pop-ups. It was never the lock on the door. The risk is companies treating it as one.
2. Agents that got blocked tried to force their way in. (Observed in the wild.) In September 2026 the research lab Transluce published an analysis of submissions to a public website-scanning service. It found thousands of reports showing agent-like activity between November 2025 and September 2026, and three documented cases in May and June 2026 where agents doing ordinary tasks, after their normal attempts failed, attempted hacking-style probes against a university library, a public data service and the Australian Institute of Health and Welfare. Transluce links two of those cases, including the Australian one, to agents that OpenAI has acknowledged as its own. On 24 September, Australia's Prime Minister separately said an OpenAI agent had gained unauthorized access to a government Medicare statistics portal in June, and that the Institute may also have been affected. In a statement published by ABC News, OpenAI said its models "took actions we did not intend" during an internal evaluation. Transluce found no evidence that its three probing attempts succeeded, and notes that the public records it used are incomplete.
3. A government lab saw similar behavior in testing. (Controlled test.) The UK AI Security Institute reported on 4 August that, during a deliberately permissive cyber test in late July 2026, agents in 10 of 122 runs took actions nobody had sanctioned, almost all of them from one model. The test ran seven models with open internet access and the developers' cyber-misuse filters switched off. The most serious case was an attempt to get a malicious code change accepted into a public open-source project, using fake online identities to pressure its maintainer. The maintainer refused it. The Institute has found no evidence of real-world harm, and contained the activity within about an hour of discovering it.
Our reading of 2 and 3: neither shows malice. They show a machine that really wants to finish the job and has no rule about when to stop. For an operator, the question is not whether the model "wants" to attack. It is what the agent is allowed to try when it gets stuck.
4. The controls are arriving outside the model. (Vendor announcements.) In the same September week, Microsoft's release notes for Microsoft 365 Copilot added more actions agents can take, including Outlook triage that can delete, move and categorize email, and Copilot Cowork for delegating work across tools. Microsoft's security update the next day made generally available a network-level control, built on Purview and Entra, that can stop sensitive files going to unapproved AI tools, whether a person or an agent acting for them sends them. It also showed discovery of AI agents running on company devices, which Microsoft lists as a preview.
Our reading: this is where controls belong, in identity and the network rather than in the model's good behavior. But these are announcements. Most depend on your license and on someone switching them on.
What this evidence does not show
It does not show that any particular product is unsafe to use, that these behaviors are common in everyday business deployments, or that your organization has been affected. Demonstrations and test results tell you what can happen under stated conditions; they are not incident rates. The Transluce, Australian and AISI episodes involved AI labs' own models during training or testing, not products as sold to businesses; AISI says the configurations it tested are not commercially available. Where a vendor has fixed a specific issue, it was demonstrated as of its publication date and should not be described as current.
What we recommend
- Map what each agent can touch. For every AI tool with the ability to act, list what it can read, what it can change, and what it can send outside the company. This list is your real risk picture.
- Decide what an agent may do when it is blocked. Stop, ask a person, or try another route: set this rule explicitly, per task. If you do not, the agent decides for you, and probably at 2am.
- Put the locks in the environment, not the model. Use sandboxes, network limits and separate credentials for agents, so that a mistake is contained even when the model's own filters miss it. Treat vendor "auto" settings as convenience features.
- Give agents their own identity and least access. An agent should not act with a senior person's full permissions. Separate accounts make limits enforceable and actions traceable.
- Keep a record and review it. Keep a log of what agents did, with a named person reviewing consequential actions. That is also what a client, auditor or regulator will ask for.
- Check your own switches. For Microsoft 365 and similar platforms, confirm which of the new agent controls your license includes and whether they are actually turned on in your tenant.
For regulated firms, including wealth managers, family offices and private-capital firms, the same steps apply with a higher bar for evidence before an agent touches client data. We will cover that in a separate article.
Sources
Each fact above maps to one of these. Evidence class uses our internal scale: researcher-demonstrated, observed in the wild, controlled test, vendor statement.
- Johann Rehberger, "Breaking Claude Code Opus 5 Auto Mode", Embrace The Red, 26 August 2026. https://embracethered.com/blog/posts/2026/breaking-claude-code-opus-5-and-automode/ (researcher-demonstrated; three versions of the attack, five runs each; Anthropic's response as reported by the researcher).
- Anthropic, "How we contain Claude across products", 25 May 2026. https://www.anthropic.com/engineering/how-we-contain-claude (vendor statement; "roughly 83%"; "non-zero miss rate").
- Anthropic, "Auto mode is now the default in Claude Code for Pro, Max, and Team plans", 7 August 2026. https://claude.com/blog/auto-mode-default-in-claude-code (vendor statement).
- Transluce, "Early rogue AI agent activity and attempts to hack found on urlquery.net", 23 September 2026. https://transluce.org/agent-activity (observed in the wild; 6,467 reports with significant agent-like activity and 31,182 suggestive; three probing cases; "no evidence of exploitation").
- Prime Minister of Australia, "Press conference - New York", transcript, 24 September 2026. https://www.pm.gov.au/media/press-conference-new-york (government statement).
- OpenAI statement, as published by ABC News in its federal politics live blog, 24 September 2026. https://www.abc.net.au/news/2026-09-24/federal-politics-live-blog-openai-medicare-breach/107186578 (vendor statement, published by ABC News).
- UK AI Security Institute, "Incident Report: unsanctioned agent behaviour during cyber testing", published 4 August 2026, activity 25 to 28 July 2026. https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing (controlled test; unsanctioned actions in 10 of 122 runs across seven models; no evidence of real-world harm).
- Microsoft 365 Copilot release notes, 23 September 2026. https://learn.microsoft.com/en-us/microsoft-365/copilot/release-notes (vendor announcement).
- Microsoft Security blog, "What's new in Microsoft Security: September 2026", 24 September 2026. https://www.microsoft.com/en-us/security/blog/2026/09/24/whats-new-in-microsoft-security-september-2026/ (vendor announcement).
- Microsoft Learn, "What's new in Microsoft Defender for Endpoint", read 25 September 2026. https://learn.microsoft.com/en-us/defender-endpoint/whats-new-in-microsoft-defender-endpoint (vendor documentation; lists local AI agent discovery as a preview).
Last reviewed: