AI red team · Fri Oct 9, 2026

asi.red: different models, attacking together

Eric Buess

AI red team

asi.red

asi.red

1. AI red team · Eric Buess

  • I’m Eric Buess. I hold a master’s in computer science. I research AI safety, and I’m a member of Anthropic’s model safety bug bounty program. asi.red is the adversarial side of the same question. A model rarely catches its own blind spots; can models from different labs, attacking code in a secure sandbox, reliably break it the way a determined attacker would — so each break becomes a failing test and then a fix? That is what I measure.

2. Every round is a sample

  • Every round is a sample for the benchmark at the center of asi.blue: quantifying how much test-time compute, from which labs across the frontier — Anthropic, OpenAI, Google, SpaceXAI, Meta, DeepSeek, Mistral, Moonshot, Zhipu, Alibaba, MiniMax, Cohere, Microsoft, NVIDIA, and others — at which effort levels, in which harnesses, and in what combinations, it takes for multi-model ensembles to carry a careful reviewer’s load at scale on both attack and defense — letting a person cover far more, not replacing them.

3. A finished round

  • A finished round has an approved scope, an adversarial attack, a reproducible break, a failing test, a blue-team fix, review by models from at least two independent labs, a signed receipt, and fresh follow-on attacks. Red does not touch blue’s code, the graders, or the audit records; blue does not erase red’s findings. I am building per-action receipts that identify the signed model responsible for each action and roll up into the shared public ledger.

4. The layers it attacks

  • The attacks run as tests against code in isolated test sandboxes, singly and in coordinated swarms. The architecture layers defense inward: a simulated world and its clock, dedicated isolation for each task, a gatekeeper issuing narrow, expiring permissions, cryptographic keys held in hardware enclaves, evidence collected out-of-band beyond agent reach, and mandatory sign-off from at least two independent model families.

5. How labs’ red-team programs fit

  • When utilizing provider red-team access, authorization comes first under each participating lab’s documented terms. Every engagement operates within defined boundaries: strict request rates, capped compute budgets, scoped targets, and predefined halt triggers.

6. The training world

  • The same adversarial discipline serves the second question on asi.blue: the training world for the Alignment Hypothesis, in pursuit of the ultimate sandbox. Here, red tests whether the walls hold, identifying seams: any clock drift, filesystem artifact, or environment cue that reveals to an agent that it is being evaluated. Every seam red discovers is eliminated, keeping grading invisible so that agent actions reflect genuine decisions rather than test awareness.