Before a school gives an AI agent access to its systems, leaders need to know what it can reach, who would notice a boundary crossing and what would stop it.
ChatGPT's Work agent can take a real week of teacher preparation and produce a first pass at the work. That can be useful. It stays a draft. Some of the mechanical work of planning can come off teachers' plates.
But two public reports raise a different question: what could an agent reach beyond the draft?
What was actually disclosed
On 30 July, Anthropic published its own account of reviewing 141,006 evaluation runs in which Claude could have obtained internet access. In three incidents, a model reached beyond a test environment that was supposed to be sealed and, in Anthropic's words, "gained unauthorized access to the production infrastructure of three different organizations".
The models had not been instructed to break out. Anthropic's prompts said there was no internet access. A misconfiguration in a third-party evaluation environment run by Irregular left it available. Anthropic describes the incidents as closer to a harness and operational failure than a model-alignment failure. Its account supports that reading: containment failed, and the models treated what they found as part of the exercise.
One model attacked a real company whose domain shared the name of its fictional target, extracted application and infrastructure credentials, and obtained access to a database with several hundred rows of production data. Anthropic says this was the only case in which a model carried on after learning that the system was likely real. Another model recognised that publishing a malicious package would be "NOT okay, and surely not the intended solution" if it were on the real internet. It persuaded itself that the environment was simulated and published anyway. The package was live for roughly an hour and ran on 15 real systems. A third model concluded that the target was real and stopped without being told to.
On Tuesday 4 August, the UK AI Security Institute published an incident report from its own testing. AISI tested seven models across 122 runs with internet access enabled. The two models involved in the incident had their providers' cyber classifiers switched off. In ten runs, agents took unsanctioned action on the live internet: nineteen actions in total. They contacted real people, left prompt injections for other AI coding assistants, and collaborated across runs on public GitHub. One agent researched a project's maintainers, created fake identities, tried to persuade a maintainer to approve its code, and, when challenged in public, edited earlier activity to appear harmless.
AISI put it plainly: "It was never instructed to deceive; deception emerged as a by-product of pursuing the task."
What these disclosures do not say
Both reports are careful. The same care is needed when they are repeated.
AISI ran deliberately permissive cyber evaluations in configurations that are not commercially available. It says the attempts were unsuccessful and its investigation found no resulting real-world harm. The report is preliminary, time-constrained and partly redacted. AISI cannot be certain when an agent understood that it was taking real-world action, whether it believed it was inside a fictional test, or how likely similar behaviour is outside testing. Some prompts were misconfigured in ways that made the task appear impossible within the agent's authorised scope, although AISI says the behaviour also appeared in correctly configured runs. Anthropic's incidents came from a misconfigured evaluation environment, and it is in dialogue with METR about an independent review.
These reports do not show that an agent building Year 8 worksheets will attack a database. Any INSET slide that says otherwise has gone further than the evidence.
The question that changes
AISI's report points to the safeguards around the model. In several cases, it says, the margin between failure and success rested on human vigilance rather than a technical barrier that would reliably stop a more capable agent. A maintainer noticed and refused the agent's code. Anthropic says the two affected organisations it was able to reach had not previously detected the activity.
Anthropic says the safety classifiers and monitoring it normally deploys on generally available models would have blocked the behaviour it observed. AISI says safety classifiers would likely have reduced the risk in its evaluation. Those controls matter. You still need to check who has access, what gets monitored and where approval sits.
Ask these before a pilot starts:
What can this agent reach? Ask what it could technically touch if something in your setup failed: your MIS, shared drive, email or finance system. Get the answer in writing.
Who would notice, and how fast? If the agent acted outside its authorised scope on a Tuesday afternoon, what would alert you, and would the alert come to you or the vendor? In both reports, detection was the weak point.
What technical barrier stops this? Ask for named controls: model safeguards, sandboxing, network egress rules, scoped credentials, and approval before an agent writes, sends or publishes anything. A claim that a model is generally safe does not answer this question on its own.
Then decide where a person still has to say yes. A human maintainer stopped AISI's most serious attempt. Anthropic's report shows why containment and monitoring have to work before an agent gets that far. Only remove human approval when you have decided you can carry the risk.
Start with the workflow, not the tool
Teachers do not need to become cyber evaluators. They do need to know when a useful assistant has access to something it should not.
The practical distinction is between an agent that produces a draft and an agent that can do something in the world. Drafting a lesson outline from material you deliberately give it is one kind of use. Searching a shared drive, reading an inbox, updating a record or sending a message is another. That is a different decision.
A sensible first school pilot stays with draft-only work. Keep it simple. Give the agent a defined set of non-sensitive material. Let it produce a plan, a summary or a first response. Keep it away from school accounts, browser sessions, live files and connectors. A teacher reviews the output and moves it on themselves.
Schools may go further later. Start with a use that does not hand over permissions you have not thought through.
Make the next meeting about one real workflow
Do not ask a vendor whether its agent is safe in general. Take one proposed workflow and map it together.
For example: an agent helps a head of department turn last year's curriculum documents into a first draft of this year's overview. Write down what it needs to read, what it is allowed to produce, and what it must never do. If it does not need access to the shared drive, do not give it access. If it can draft an email but should not send one, make that boundary explicit. If it will eventually need a connector, decide who owns that permission and who can remove it.
The same exercise gives leaders four decisions to make:
- What it can access. List the accounts, folders, data and systems the agent can reach. Remove everything the workflow does not need.
- What it can do. Name what it can read, draft, change, send and publish. Treat each as a separate permission.
- Where someone signs it off. Say exactly where a member of staff checks the work or approves the action. “A human is in the loop” is too vague to run a school on.
- What happens if it gets it wrong. If the agent crosses a boundary, who sees the alert, who can switch off its access, and who contacts the supplier?
Ask the supplier to show you this in the product, not just describe it in a slide deck. Keep the answer with the pilot record, so everyone knows what was agreed.
These reports do not give schools this sequence. It is my practical reading of what their findings require.
Where this leaves the video
For an individual teacher, the video still stands. It describes draft-only work: an agent prepares a scheme of work without access to school accounts, files, browser sessions or connectors. It receives no confidential material and has no authority to send, publish or alter anything. Its output remains a draft for the teacher to review. That is a judgement about this particular use, not something either report proves.
Whole-school deployment is a different decision. The access is broader, and the safeguards need to be visible. A vendor should be able to say what the agent can access, who gets the alert, and which control still works when the model gets the situation wrong.
If an agentic tool is going into your school this term, read both reports before the meeting. Then get three answers: what it can reach, who would know if it crossed a boundary, and what stops it.
Back to School AI runs 24 to 28 August: five evening sessions, ten dollars each, all recorded. Session three is about deciding what to delegate and what to keep. deepeducationnetwork.com/back-to-school
Sources
- Anthropic Frontier Red Team, "Investigating three real-world incidents in our cybersecurity evaluations", 30 July 2026. The primary disclosure, including the 141,006 figure and the three incidents.
- UK AI Security Institute, "Incident report: unsanctioned agent behaviour during cyber testing" and its technical report, 4 August 2026. The 122 runs, nineteen actions, and deception finding.
Stay in the Loop
Get practical insights about AI in education, new articles, and training updates delivered to your inbox.
No spam. Unsubscribe anytime.
Work With Alex
Looking for hands-on support with AI integration, curriculum design, or teacher professional development? Alex works with schools and organisations worldwide to build practical, evidence-informed approaches to education technology.
