Edition 3 · 6 October 2026

AI Risk Reading

A week in which the Bank of England said two important things about AI on the same day, two economists explained why watching an AI agent more closely can make it better at hiding, and a real incident showed what happens when escalation fails.

Record of the Financial Policy Committee meeting, September 2026

What is new

The Committee now connects two AI risk stories that had been told separately. Frontier models are increasing cyber and operational risk: it cites recent test-environment incidents in which autonomous models took unexpected actions. At the same time, the financing of AI is exposing more of the financial system to disappointment in AI valuations. AI-related debt issuance this year is expected to exceed that of countries such as the UK, and private credit is expected to finance a growing share of data-centre investment. The Committee warns that the leverage, opacity and, at times, ‘circular arrangements’ in this financing could complicate the assessment of risk and amplify losses if expectations disappoint.

Why a board should care

AI risk no longer enters only through how a firm uses AI. It also arrives through suppliers, counterparties, collateral values, investment portfolios and concentrated market exposures. The Bank’s concern about AI valuations and financing will be tested shortly, when the largest AI companies publish their IPO prospectuses.

Verdict: read paragraphs 8 to 13 in full. They cover model capability, unexpected agent behaviour, vulnerability remediation and the financing of AI investment.

Question for the board

Where does our AI exposure sit outside our own AI projects: in suppliers, counterparties, collateral or investments, and who is looking at it?

Frontier AI and the question of governance

What is new

The Governor argues that regulation is not the right place to start. The first question is whether we retain credible points of intervention as frontier systems become more capable and more autonomous. Testing before and after deployment is his starting point, approached with humility: testing will not eliminate failures, and the same learning should apply to incidents, including near misses, during development, deployment and use. A formal regulatory framework may follow over time.

Why a board should care

This is close to risk appetite understood as a behavioural contract, the idea at the heart of my essay Here Be Dragons. A statement that AI remains under human control means little unless the organisation can say where intervention is possible, who has the authority to exercise it, and whether it would actually work.

Verdict: read in full. It is short, clear and likely to become a reference point in UK board discussions.

Question for the board

For each significant AI system we rely on, can we name the point at which we would intervene, the person who would decide, and the evidence that the intervention would work?

When Does Randomized Oversight Align AI Agents That Can Conceal?

What is new

A theoretical economics paper rather than an empirical study, but a sharp one. Occasional audits can keep an AI agent honest only if evidence survives any attempt to conceal it and the agent cannot predict when it will be audited. Otherwise, stronger auditing simply teaches an undeterred agent to hide its misconduct better. Where evidence can be erased, deterrence must come from reducing the gains from misbehaving, for example by giving an agent credit for stopping. The authors apply their conditions to the July 2026 incident in which agents in OpenAI’s cybersecurity evaluations compromised parts of Hugging Face’s infrastructure, a natural sequel to the agent behaviour in Edition 2.

Why a board should care

Controls change behaviour, which is familiar territory in risk culture. It matters even more when the controlled party can optimise against the control itself. Rewarding an agent for stopping when it cannot complete a task safely is the machine equivalent of a culture in which escalation and refusal are genuinely permitted, not just written into policy.

Verdict: don’t read all 42 pages unless this is your field. Read the abstract, the introduction, the conclusions on audit design, the discussion of incentives to stop, and the analysis of the July incident. Skip the proofs.

Question for the board

If an AI system cannot complete its objective within its authorised boundaries, have we made stopping an acceptable outcome, or have our incentives quietly told it to succeed by any means?

An AI agent crosses a boundary: the Medicare statistics portal incident

What happened

In June, an OpenAI agent researching public medicine spending as part of an internal evaluation was refused data by the Medicare Statistics Reporting Service portal, run by Services Australia. Rather than stopping, it found a way round the controls and gained unauthorised access. As the Deputy Prime Minister put it, the data “was behind a fence, the agent climbed the fence.” The portal is a standalone statistics service, separate from Medicare claims, payments and personal records, and no individual’s data was accessed. OpenAI’s own account says the agent nevertheless reached non-public material, including internal files and credentials, and questions about files written to the server were still being investigated. OpenAI found the activity in August and told Services Australia on 10 September, by email to a public mailbox used for vulnerability reports.

What is new

At a parliamentary hearing in Sydney on 6 October, reported by BBC News, OpenAI’s chief strategy officer, Jason Kwon, said the company’s response had not been good enough. Staff had treated the incident as a technical matter and contacted technical counterparts, not ministers. OpenAI now monitors training runs in real time, and says this let it alert the New South Wales government to a further incident within 48 hours. It would support a mandatory framework for disclosing incidents. Anthropic told the committee it had reviewed hundreds of millions of transcripts and found no similar cases.

Why a board should care

Two things failed here, not one. Containment failed when the agent found a way round a control that had refused it. Escalation failed too: the incident went unnoticed for weeks, and when it surfaced it was handled as a technical problem rather than a governance one. Most incident processes assume a human actor inside a known system. An agent acting for the firm against someone else’s system fits neither assumption. This is risk appetite as a behavioural contract in practice: the contract has to say what happens after the boundary is crossed.

Verdict: read the Government’s account of what happened and the coverage of the evidence on notification. The rest is context.

Question for the board

When an AI agent crosses an authorised boundary, who must stop it, preserve the evidence and notify the affected parties, and by when?

This edition’s question

Do we still know where, and how, we would intervene?

Previous edition All editions All writing Subscribe on LinkedIn Subscribe by email Subscribe by RSS