AI red team · Fri Oct 9, 2026

asi.red: the red side of asi.blue

Eric Buess

AI red team

asi.red

asi.red

1. AI red team · Eric Buess

  • asi.red is the red side of asi.blue. Models from different labs attack my own code — never anyone else’s — so that a break becomes a failing test and then a fix. One lab’s model attacks; another lab’s model writes the fix. Difference is the point: a model unlike the one that wrote the code is the one most likely to break it.

2. Every round is a sample

  • Every round is a sample for the benchmark at the center of asi.blue: how much test-time compute, from which frontier and near-frontier families — Claude, GPT, Gemini, Grok, DeepSeek and others — in what combinations, it takes for models working in concert to match or beat a careful human reviewer, on attack and on defense.

3. A finished round

  • A finished round has an approved scope, a real attack, a break that reproduces, a failing test, a blue-team fix, review by models from two other labs, a signed receipt, and fresh attacks. Red does not touch blue’s code, the graders or the records; blue does not erase red’s findings. Those receipts are being built to name the signed model that did the work and to roll up into the same public ledger as asi.blue.

4. The layers it attacks

  • The attacks run as tests against my own code in my own testing labs, and the design layers the targets inward — a simulated world and its clock, a sandbox for each job, a gatekeeper with narrow, expiring permission, keys held in an enclave, evidence collected outside the agents, and review by two other families.

5. How labs’ red-team programs fit

  • When I use a lab’s red-team program, approval comes first — from me and from each lab whose models are used, under that lab’s own terms. Every engagement has a scope, caps on attackers, requests and spend, and stop conditions. If a lab says no, that work stops; it never moves quietly to another lab’s access.

6. The training world

  • The same discipline serves the second question on asi.blue, the training world for the Alignment Hypothesis. There, red’s job is to find the seams: any clock, file or quirk that tells an agent it’s being tested. Every seam red finds gets closed, so the grading stays invisible and an agent’s choices mean what they seem to.