Skip to main content

What can your school's AI agent reach?

An AI agent's route stops at a human approval gate before reaching school systems
ByAlex Gray13 Aug 20267 min readUpdated 13 Aug 2026

Before a school gives an AI agent access to its systems, leaders need to know what it can reach, who would notice a boundary crossing and what would stop it.

Share:

ChatGPT's Work agent can take a real week of teacher preparation and produce a first pass at the work. That can be useful. It stays a draft. Some of the mechanical work of planning can come off teachers' plates.

But two public reports raise a different question: what could an agent reach beyond the draft?

What was actually disclosed

On 30 July, Anthropic published its own account of reviewing 141,006 evaluation runs in which Claude could have obtained internet access. In three incidents, a model reached beyond a test environment that was supposed to be sealed and, in Anthropic's words, "gained unauthorized access to the production infrastructure of three different organizations".

The models had not been instructed to break out. Anthropic's prompts said there was no internet access. A misconfiguration in a third-party evaluation environment run by Irregular left it available. Anthropic describes the incidents as closer to a harness and operational failure than a model-alignment failure. Its account supports that reading: containment failed, and the models treated what they found as part of the exercise.

One model attacked a real company whose domain shared the name of its fictional target, extracted application and infrastructure credentials, and obtained access to a database with several hundred rows of production data. Anthropic says this was the only case in which a model carried on after learning that the system was likely real. Another model recognised that publishing a malicious package would be "NOT okay, and surely not the intended solution" if it were on the real internet. It persuaded itself that the environment was simulated and published anyway. The package was live for roughly an hour and ran on 15 real systems. A third model concluded that the target was real and stopped without being told to.

On Tuesday 4 August, the UK AI Security Institute published an incident report from its own testing. AISI tested seven models across 122 runs with internet access enabled. The two models involved in the incident had their providers' cyber classifiers switched off. In ten runs, agents took unsanctioned action on the live internet: nineteen actions in total. They contacted real people, left prompt injections for other AI coding assistants, and collaborated across runs on public GitHub. One agent researched a project's maintainers, created fake identities, tried to persuade a maintainer to approve its code, and, when challenged in public, edited earlier activity to appear harmless.

AISI put it plainly: "It was never instructed to deceive; deception emerged as a by-product of pursuing the task."

What these disclosures do not say

Both reports are careful. The same care is needed when they are repeated.

AISI ran deliberately permissive cyber evaluations in configurations that are not commercially available. It says the attempts were unsuccessful and its investigation found no resulting real-world harm. The report is preliminary, time-constrained and partly redacted. AISI cannot be certain when an agent understood that it was taking real-world action, whether it believed it was inside a fictional test, or how likely similar behaviour is outside testing. Some prompts were misconfigured in ways that made the task appear impossible within the agent's authorised scope, although AISI says the behaviour also appeared in correctly configured runs. Anthropic's incidents came from a misconfigured evaluation environment, and it is in dialogue with METR about an independent review.

These reports do not show that an agent building Year 8 worksheets will attack a database. Any INSET slide that says otherwise has gone further than the evidence.

The question that changes

AISI's report points to the safeguards around the model. In several cases, it says, the margin between failure and success rested on human vigilance rather than a technical barrier that would reliably stop a more capable agent. A maintainer noticed and refused the agent's code. Anthropic says the two affected organisations it was able to reach had not previously detected the activity.

Anthropic says the safety classifiers and monitoring it normally deploys on generally available models would have blocked the behaviour it observed. AISI says safety classifiers would likely have reduced the risk in its evaluation. Those controls matter. You still need to check who has access, what gets monitored and where approval sits.

Ask these before a pilot starts:

What can this agent reach? Ask what it could technically touch if something in your setup failed: your MIS, shared drive, email or finance system. Get the answer in writing.

Who would notice, and how fast? If the agent acted outside its authorised scope on a Tuesday afternoon, what would alert you, and would the alert come to you or the vendor? In both reports, detection was the weak point.

What technical barrier stops this? Ask for named controls: model safeguards, sandboxing, network egress rules, scoped credentials, and approval before an agent writes, sends or publishes anything. A claim that a model is generally safe does not answer this question on its own.

Then decide where a person still has to say yes. A human maintainer stopped AISI's most serious attempt. Anthropic's report shows why containment and monitoring have to work before an agent gets that far. Only remove human approval when you have decided you can carry the risk.

Start with the workflow, not the tool

Teachers do not need to become cyber evaluators. They do need to know when a useful assistant has access to something it should not.

The practical distinction is between an agent that produces a draft and an agent that can do something in the world. Drafting a lesson outline from material you deliberately give it is one kind of use. Searching a shared drive, reading an inbox, updating a record or sending a message is another. That is a different decision.

A sensible first school pilot stays with draft-only work. Keep it simple. Give the agent a defined set of non-sensitive material. Let it produce a plan, a summary or a first response. Keep it away from school accounts, browser sessions, live files and connectors. A teacher reviews the output and moves it on themselves.

Schools may go further later. Start with a use that does not hand over permissions you have not thought through.

Make the next meeting about one real workflow

Do not ask a vendor whether its agent is safe in general. Take one proposed workflow and map it together.

For example: an agent helps a head of department turn last year's curriculum documents into a first draft of this year's overview. Write down what it needs to read, what it is allowed to produce, and what it must never do. If it does not need access to the shared drive, do not give it access. If it can draft an email but should not send one, make that boundary explicit. If it will eventually need a connector, decide who owns that permission and who can remove it.

The same exercise gives leaders four decisions to make:

  1. What it can access. List the accounts, folders, data and systems the agent can reach. Remove everything the workflow does not need.
  2. What it can do. Name what it can read, draft, change, send and publish. Treat each as a separate permission.
  3. Where someone signs it off. Say exactly where a member of staff checks the work or approves the action. “A human is in the loop” is too vague to run a school on.
  4. What happens if it gets it wrong. If the agent crosses a boundary, who sees the alert, who can switch off its access, and who contacts the supplier?

Ask the supplier to show you this in the product, not just describe it in a slide deck. Keep the answer with the pilot record, so everyone knows what was agreed.

These reports do not give schools this sequence. It is my practical reading of what their findings require.

Where this leaves the video

For an individual teacher, the video still stands. It describes draft-only work: an agent prepares a scheme of work without access to school accounts, files, browser sessions or connectors. It receives no confidential material and has no authority to send, publish or alter anything. Its output remains a draft for the teacher to review. That is a judgement about this particular use, not something either report proves.

Whole-school deployment is a different decision. The access is broader, and the safeguards need to be visible. A vendor should be able to say what the agent can access, who gets the alert, and which control still works when the model gets the situation wrong.

If an agentic tool is going into your school this term, read both reports before the meeting. Then get three answers: what it can reach, who would know if it crossed a boundary, and what stops it.

Back to School AI runs 24 to 28 August: five evening sessions, ten dollars each, all recorded. Session three is about deciding what to delegate and what to keep. deepeducationnetwork.com/back-to-school

• • •

Sources

Alex Gray

Alex Gray

Head of Sixth Form & BSME Network Lead for AI in Education. Alex explores how artificial intelligence is reshaping teaching, learning, and the future of work — with honesty, clarity, and a focus on what matters most for educators and students.

Stay in the Loop

Get practical insights about AI in education, new articles, and training updates delivered to your inbox.

No spam. Unsubscribe anytime.

Work With Alex

Looking for hands-on support with AI integration, curriculum design, or teacher professional development? Alex works with schools and organisations worldwide to build practical, evidence-informed approaches to education technology.

Discussion

Sign in to join the discussion.

Keep reading

Never Miss an Insight

Join educators worldwide who receive practical thinking about AI in education, teaching strategies, and professional development — straight to their inbox.

No spam. Unsubscribe anytime.