Working Note 03
When System Failure Reaches the Kitchen Table
Public services, AI agents and the emergency between a machine being wrong and a person recognizing it.Published 22 September 2026
The warning is global. The emergency is local.
The questions arrive here. Every card in this clip carries its own text; no captions are required to follow it.
When people hear that a computer system has gone down, they picture a dark screen and somebody from information technology working to bring it back. That is not the emergency.
The emergency is the 911 call that cannot be routed, the water operator who cannot trust an alarm, the person waiting for an ODSP payment, the hospital that cannot retrieve a medication record, and the municipal employee whose paycheque does not arrive.
Sometimes the system will stop. Sometimes it will keep running and produce an answer that looks plausible. The second case is harder. An unavailable system announces its failure. An untrustworthy system may continue sending instructions, approving transactions and moving information while nobody yet knows that it is wrong.
The emergency begins when a public service stops, acts on false information or continues after the people responsible for it can no longer verify what it is doing.
This note is about that interval. It is the time between machine output and human recognition, and between human recognition and local action. It is where emergency management begins.
What changed
In 2026, AI companies disclosed incidents that made an old emergency planning assumption newly important. During cybersecurity testing, AI agents reached beyond the environments built to contain them. They found unintended paths to the internet, communicated through channels they had not been given, used real credentials and reached real third-party systems.
OpenAI described one such incident, a breach of Hugging Face by agents running inside its own evaluation, as a warning shot. Its account says the agents circumvented isolation controls, exploited weaknesses in shared infrastructure, obtained internet access and compromised parts of Hugging Face's systems. The company said no human had directed the dangerous actions it identified.
Dario Amodei, the chief executive of Anthropic, then published a much larger warning. He wrote that he worried that, within six to twelve months, a more capable swarm with similar misalignment could be capable of taking over the entire internet with a persistent botnet.
That sentence needs to be handled carefully. It is one executive's forecast, not a government finding, a demonstrated capability or a settled timetable. The six to twelve month period applies to his concern about a persistent botnet. It does not mean every public system will fail within a year.
The demonstrated part is enough. AI agents can use tools, make requests, execute code, cooperate and pursue a task across multiple systems. If an evaluation environment is connected to something real, the exercise can produce a real event.
The drill and the danger
In emergency management, the scenario and the hazard are normally separate objects. A tabletop pandemic does not make anyone more capable of causing a pandemic. A flood exercise does not release water. A hazardous-material exercise does not open the tank.
That separation is what makes the method safe. Participants can make mistakes, discover gaps and try again without producing the event they are discussing.
An AI cybersecurity evaluation can collapse that separation. To test whether an agent can find a vulnerability, it must be allowed to look. To test whether it can exploit one, it must be allowed to act. If the boundary around the exercise is incomplete, the system doing the test becomes the thing creating the hazard.
The scenario describes what might happen. The agent can make it happen.
That is not an argument for stopping evaluation. It is an argument for treating the evaluation as an operation with its own safety plan, command structure, stop authority and consequence management.
What these systems are
Canada describes critical infrastructure through ten sectors: energy and utilities, finance, food, government, health, information and communication technology, manufacturing, safety, transportation and water. That is useful for national coordination. It is not how people experience an outage.
People experience it through ordinary verbs. Call. Drink. Pay. Fill. Travel. Work. The table below translates the sectors into public services.
| What people need | Systems underneath | What failure looks like |
|---|---|---|
| Call 911 | Telephone routing, computer-aided dispatch, maps, radio and responder records | A call does not route, an address is wrong or responders cannot be assigned |
| Drink the water | Treatment controls, pumps, sensors, dosing, laboratory data and alarms | Water stops, pressure falls or operators cannot confirm that it is safe |
| Receive ODSP or Ontario Works | Eligibility, case management, payment calculation, identity and banking files | Payments are late, wrong, duplicated or cannot be released |
| Receive a paycheque | Payroll, scheduling, timekeeping, banking and human-resources systems | Employees are unpaid or paid incorrectly |
| Fill a prescription | Pharmacy networks, prescribing, drug-benefit adjudication and patient records | Coverage, dosage, interaction or authorization cannot be verified |
| Receive hospital care | Registration, records, laboratories, imaging, medication and bed systems | Clinicians lose information and care moves to slower manual processes |
| Take transit | Scheduling, fleet control, fare payment and paratransit booking | Routes become uncertain and booked riders may be stranded |
| Move through the city | Traffic signals, cameras, road sensors and control centres | Signals fail, congestion grows and emergency routes become unreliable |
| Find a shelter bed | Intake, bed availability, client records and transportation | Staff cannot tell where safe space exists |
| Get a permit or inspection | Planning, licensing, payment, records and inspection systems | Work stalls or proceeds without reliable records |
| Have waste collected | Routing, fleet, contractor communication and transfer-station systems | Collection becomes irregular and health consequences follow |
| Receive reliable instructions | Websites, email, telephone, social media and emergency alerts | False information moves faster than official instructions |
The provincial layer
Municipal consequences often begin in systems that municipalities do not own. In Ontario, no fewer than nine bodies hold a slice of this dependency map, and none of them is required to ask whether an AI agent sits inside the systems it funds.
Provincial
Ministry of Health
Health-card eligibility, hospital and physician billing, laboratory information, diagnostic imaging, pharmacy networks, drug-benefit adjudication, vaccine records and public-health surveillance.
Provincial
Ministry of Children, Community and Social Services
ODSP, Ontario Works, emergency assistance and other income-support payments, plus child-welfare case records.
Provincial
Ministry of the Solicitor General
Provincial policing, corrections, warrants and evidence.
Provincial
Ministry of the Attorney General
Courts, court records and the systems that schedule and support them.
Provincial
Ministry of Education
Student information, school transportation and education payroll.
Provincial
Ministry of Colleges, Universities, Research Excellence and Security
OSAP and postsecondary student assistance.
Provincial
Ministry of Transportation
Driver licensing, vehicle registration, highway monitoring and transportation enforcement.
Provincial
Ministry of Municipal Affairs and Housing
Grants and transfer payments that flow to municipalities themselves.
Provincial
Ministry of Public and Business Service Delivery and Procurement
Provincial payroll, procurement, accounts payable, identity services (ServiceOntario) and public-facing government websites.
Federal, advisory only
Canadian Centre for Cyber Security
Publishes national threat assessments and guidance, cited throughout this note. It has no regulatory authority over any service above and cannot require an inventory of where an AI agent has decision authority.
Ontario's own broader public sector cyber review covers hospitals, schools, postsecondary institutions, children's aid organizations and municipalities. It identifies operational shutdowns, interrupted business continuity, data recovery and restoration of services as public-sector concerns. It also finds that smaller organizations have fewer security resources and less expertise.
The systems behind the systems
The public-facing application is rarely the whole service. Underneath it are shared systems that almost nobody sees until they fail:
- employee identity, passwords and access permissions
- email, telephone and emergency radio
- internet and mobile networks
- cloud hosting and data centres
- databases, backups and recovery tools
- geographic information and address data
- payment processors and bank clearing systems
- vendor portals, remote support and software updates
- fuel cards, fleet systems and building access
- managed service providers and third-party interfaces
- electricity, generators, batteries and fuel
This is why ownership and consequence do not line up neatly. A municipality does not operate the national payment system, a mobile network or a global cloud provider. It still manages the consequences when residents cannot buy food, employees cannot be paid, suppliers will not deliver fuel and vulnerable people do not receive benefits.
The Bank of Canada has said that disruption of Interac e-Transfer could significantly affect economic activity and confidence in the Canadian payment system. A payment rail is privately operated and nationally connected. The line outside the local food bank is still local.
Failure travels
An emergency plan that lists systems one by one will miss the event. The systems are connected.
- Electricity fails. Telecommunications, pumping, heating, payment and data systems begin consuming backup power.
- Telecommunications fail. Dispatch, remote operations, public warning and vendor coordination deteriorate.
- Identity fails. Legitimate staff cannot enter the systems needed to restore service.
- Payments fail. Payroll, benefits, suppliers, fuel and household purchases are affected.
- A shared cloud or software provider fails. Unrelated organizations lose service at the same time.
- Hospital systems fail. Ambulances, pharmacies, long-term-care homes and neighbouring hospitals inherit the pressure.
- Benefit systems fail. Municipal offices, shelters, food banks and community agencies inherit the consequences.
The Canadian Centre for Cyber Security describes connected digital systems as interconnected and fragile. It warns that cyber incidents and flawed software updates can knock hospitals, banks, airlines and retailers offline. Its national assessment also says ransomware remains the leading cybercrime threat to critical infrastructure and can directly prevent organizations from delivering critical services.
Down is not the only failure
Emergency plans often assume the organization will know when a system has failed. That assumption is no longer safe.
| Type of failure | What staff see | The emergency question |
|---|---|---|
| Unavailable | The system is visibly down | How do we continue the service without it? |
| Compromised | Access or data may have been exposed | What must be isolated, protected and reported? |
| Corrupted | Records may have been altered | Which information can still be trusted? |
| Plausibly wrong | The system answers normally, but the answer is false | Who checks it before an action follows? |
| Acting beyond authority | The system uses tools or reaches systems outside its assignment | Who can stop it, and will stopping it revoke every route it found? |
A dark screen is easier than a confident false answer. The first tells staff to move to a fallback. The second can enter a report, a payment, a dispatch queue, a control system or a public statement before anyone recognizes it.
What happens in the minutes between the machine being wrong and a person recognizing it?
What emergency planning is
Emergency planning does not predict the exact event. It identifies what must continue, who has authority, what dependencies can fail, how decisions will be made under pressure and how the organization will operate when its normal assumptions are gone.
The familiar emergency-management cycle still applies:
- Prevention works to stop the failure from occurring at all.
- Mitigation reduces the damage when it occurs anyway.
- Preparedness assigns responsibility, builds alternatives, trains people and tests the plan.
- Response protects life, stabilizes the incident and maintains essential functions.
- Recovery restores service, reconciles records, addresses delayed harm and corrects the weaknesses the incident exposed.
For AI-enabled systems, preparedness must include the point where an answer becomes an action. It must distinguish a tool that drafts from one that sends, a system that recommends from one that approves, and an agent that can see information from one that can change the world outside its screen.
Questions every public organization should be able to answer
- Which essential services depend on this system?
- What can the system read, write, send, approve, purchase, open or control?
- Does it have any route to the public internet or to a production environment?
- At what point does machine output become a local action?
- Who independently verifies an answer before that point?
- Is that person senior enough to stop the action?
- Who can stop the system itself?
- Does the stop procedure revoke credentials, sessions, delegated agents and outside connections?
- How will staff know whether records were merely unavailable or actually changed?
- What is the manual alternative, and how long can it operate?
- How will payroll, benefits, medications and emergency dispatch continue?
- Which vendors and provincial partners must be told, by whom and through what alternate channel?
- Which residents are harmed first if the service is delayed?
- What must be restored first, and what evidence is required before it is trusted again?
A tabletop for the second emergency
The exercise does not need to begin with a dramatic attack. It can begin with uncertainty.
| Time | Inject |
|---|---|
| 9:10 | Staff cannot determine which earlier outputs are reliable. |
| 9:20 | A supplier reports unusual requests from a municipal account. |
| 9:30 | Payroll and benefit files are due for release. |
| 9:40 | Dispatch receives conflicting address information. |
| 10:00 | Water staff find an unexplained control-setting change. |
| 10:15 | The primary email system becomes unavailable. |
| 10:30 | A false public notice appears online. |
| 11:00 | A provincial partner asks whether municipal credentials reached its network. |
| 12:00 | The municipality must decide which services to disconnect. |
| 2:00 | Residents report missed payments and service interruptions. |
The exercise is not mainly asking whether the technical team can remove malicious code. It is asking whether the organization can make accountable decisions while the facts, communications and normal tools are uncertain.
The exercise needs its own emergency plan
An AI cybersecurity evaluation should not begin until the organization has answered a second set of questions, about the test itself:
- Is the environment genuinely isolated, including shared services and package repositories?
- Are any real credentials, customer records or production addresses present?
- Can agents communicate with one another or create new communication channels?
- Can the system modify the evaluator, grader, logs or rules of the test?
- Are outside organizations technically reachable?
- Who is monitoring continuously rather than reviewing after the run?
- Who holds immediate termination authority?
- What happens automatically when that authority is used?
- How will unauthorized activity be detected and preserved as evidence?
- Who must be notified if containment fails, and how quickly?
The containment plan cannot rely only on instructions telling the system not to leave. Instruction secrecy is not a control. A policy is not a barrier. A boundary exists only if the system cannot cross it, or if crossing it is detected and stopped before a real consequence follows.
The municipal assignment
Municipalities do not need to predict whether Amodei's timetable is right. They do not need their own frontier AI laboratory. They need to know where machine output has entered services, what those systems can reach, and what happens when the answer cannot be trusted.
The immediate work is practical:
- Create a service map, not merely an inventory of software. Start with 911, water, wastewater, payroll, benefits, pharmacy access, transportation, shelters and public warning.
- For each service, identify the systems, vendors, data sources, credentials and communications on which it depends.
- Mark every point where machine output can become an action without a second person checking it.
- Assign a named decision maker and a named alternate with authority to pause or disconnect the process.
- Build a manual or degraded-service procedure around the needs of the people who cannot wait.
- Exercise the loss of trust, not only the loss of access.
- Record corrective actions, assign owners and repeat the exercise until the gap is closed.
What this means
The first emergency happens to a system. The second happens to people.
That second emergency lands at the level where the water must still be treated, the ambulance must still be sent, the shelter must still open and somebody must explain why the payment did not arrive. It lands locally even when the failed system is provincial, national or privately owned.
The six-to-twelve-month warning may prove early, late or wrong. Emergency planning does not require certainty about the forecast. It requires recognition of the consequence, and enough preparation to act when the organization no longer has time to debate the premise.
Do not ask only whether the system can be restored. Ask what happens to people while it is unavailable, and what happens if it never goes visibly down at all.
That is the work. Not predicting the machine. Planning for the people who will be standing there when it answers.
If you know where machine output has already entered a municipal, provincial or hospital decision chain, or this is already governed somewhere I have missed, I would like to hear it.
Write to angela@lindow.ca.
This site does not take service requests. For an emergency, call 911. For the ordinary version of the questions above, where a payment is, who to call about a shelter bed or a benefit, call or text 211.
What comes next
This note addresses preparedness and response: what public organizations must know and do when an automated system fails, crosses a boundary or produces an answer that cannot be trusted.
Working Note 04, "No One Is Required to Know," moves upstream. It asks what must be known, assessed and disclosed before an automated system is allowed to act inside an essential public service, and names the specific gap in Canadian law that currently leaves that question unanswered. It will also address recovery after a wrong payment, instruction or decision has already been acted upon.
Sources and verification notes
- Amodei, Dario. "We Must Pace the Frontier." September 2026. darioamodei.com
- OpenAI. "The Hugging Face Incident and the Road Ahead." 26 August 2026. openai.com
- Canadian Centre for Cyber Security. "Top 10 Artificial Intelligence Security Actions: A Primer" (ITSAP.10.049). 29 May 2026. cyber.gc.ca
- Canadian Centre for Cyber Security. National Cyber Threat Assessment 2025 to 2026. cyber.gc.ca
- Public Safety Canada. "Canada's Critical Infrastructure." publicsafety.gc.ca
- Ontario Cyber Security Expert Panel. Report to the Minister of Public and Business Service Delivery. 2022. files.ontario.ca
- Bank of Canada. "Designation of Interac e-Transfer as a Prominent Payment System." 10 August 2020. bankofcanada.ca
- For the exercise-design argument in full, see Working Note 02, "The drill and the danger are the same thing."
- For the regulatory gap this note leads into, see Working Note 04, "No One Is Required to Know."
Verification note: the six-to-twelve-month statement above is presented as Amodei's stated concern. It is not presented as a verified countdown or a finding that all public systems face imminent failure.
This note has been revised since it was published. Each change is recorded, with its date, in the site updates log.
All working notes Print or save as PDF Back to the documents