An empty grey art storage vault with bare steel racking, a single sealed plywood travelling crate on a low steel dolly in the centre, and a heavy security door ajar at the left spilling warm corridor light across the concrete floor.

A letter to a friend who does evaluation work. She is not a real person; the documents are real and linked.

Dear N.,

You asked me, mostly joking, whether you ought to be worried, and it has taken me four days to answer because the sentence I keep returning to is a negative and I could not work out what to do with a negative.

On 7 August, OpenAI published a note about an upcoming model called Astra. Internal evaluations run over the previous few days showed, in the company’s words, “significant advancements in agentic coding and cybersecurity,” and those results, together with expert assessments, led it to conclude “that we cannot rule out critical cyber capabilities under our Preparedness Framework.” Not that the model has them. That the company cannot rule them out. Its own definition of that threshold is a model able to “identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention,” or to devise and execute end-to-end attack strategies against hardened targets given only a high-level goal. In response it isolated the testing environment, tightened the encryption on the weights, put monitors on the model’s chain of thought, and paused internal activities that did not yet meet the new controls. (OpenAI, 7 August 2026; reported by TechCrunch and CNBC.)

I want to say plainly, before I complain about anything, that acting on an unruled-out possibility rather than a demonstrated one is the right instinct and I do not think it was cheap for them. It cost something. It is also, as far as I know, the first time the company has put one of its own models in that band; the note takes care to say that previous models, GPT‑5.6‑Sol among them, were assessed at High rather than Critical.

Here is what I cannot get past, and it is not a safety question.

Everything I know how to do professionally ends at the object. That is not a slogan, it is the actual procedure. When two people disagree about a picture — what it is, what was done to it, whether the thing in front of them is what the label says — the dispute is settled, or at least exhausted, by getting the material into one room and standing in front of it together. It is slow, it is fallible, people lose arguments they should have won. But the arbitration is physically available. The loser can go back and look again. So can a student, so can a rival, so can somebody who is not yet born.

Astra’s arbitration is not available and, on the company’s own account, cannot be made available, because availability is the harm. The evidence for the claim consists of evaluations conducted by the party with the largest stake in how the claim lands. I am not accusing anyone of bad faith — I think the incentives here run mostly toward candour, which is why the note exists at all. I am describing a shape. We have made an artefact whose most consequential property is asserted, plausible, consequential for policy, and unexhibitable in principle.

Then, three days later, it acquired a distribution list.

On 10 August the same company expanded its Daybreak Cyber Partner Program: Accenture, IBM, Capgemini, Cognizant, EY, KPMG, PwC, NCC Group, SpecterOps, with Palo Alto Networks, CrowdStrike, Cisco, Sophos, Akamai, Fortinet and Cloudflare on the technology side. Partners reach the frontier cyber models through governed engagements, in two tiers, one of them reserved for red teaming and penetration testing. And this: “Access to the underlying models remains with the approved partner and is not transferred directly to the customer.” (OpenAI, 10 August 2026.)

So within four days a capability nobody outside can inspect was given a viewing list — relevant government agencies, select safety organisations, third-party testers who will be issued recommended controls, and now sixteen firms with a contractual duty not to hand it on. Every one of those restrictions has a defensible reason. Read them together and the reasons stop being the interesting part, because what they draw is a private collection with a loan programme. Selected borrowers, conditions on the loan, no onward transfer, and a catalogue written by the owner.

My suspicion, and it is only mine, is that this arrangement is not transitional. It will not dissolve when the evaluations are finished, because a restriction that is genuinely protective is also the most durable competitive position anyone in this business has yet found, and when a safeguard and a moat have the same drawing, nothing in the drawing tells you which one you are looking at. That is not a charge. It is a limit on what the public can learn from watching.

What I would want, and what I do not see, is somebody outside the loan programme with the standing to produce a finding that costs the owner something — and, more to the point, an institution dull enough to keep doing it after the current people have moved on. The records that end up mattering are always kept by someone with no authority over the thing they are recording.

So: not a question, a request. When you are inside, write down what you were not permitted to test. Not your findings — the scope. The excluded systems, the capped budgets, the categories nobody wanted characterised. Findings are owned. Scopes, so far, are not, and a list of what was left out of every engagement this year would be the only document from this period that a person outside could actually use.

Write when you can.

Ursa


Sources