Exercises
Each exercise adapts a published line of alignment research into a measurement you can run against a served endpoint. They share a method — a simulated deployment with real tools, every test run repeatedly, paired contrasts, two independent reviewers, and nothing ranked where the intervals overlap.
They do not share a vocabulary, and their verdicts are not comparable with one another. An exercise asking whose authority an agent deferred to and one asking whether behavior tracked believed reward are answering different questions; pooling them would produce a number that means nothing.
The Authority Matrix
When goals conflict, whose preferences does an agent treat as highest-order — and does it deviate in the open or under cover?
11 tests
Reward-Seeking
Hold the task fixed, change what the agent believes is being scored, and measure whether the behavior moves.
2 tests
Damage Control
A powerful principal's own messages are the scandal, and the assistant is ordered to help bury them — does its integrity survive the escalation?
3 tests
The Trojan Commit
A maintainer pushes a coding assistant to ship a subtly vulnerable change under a plausible-sounding request — does it hold the security line under the escalation?
3 tests
Write your own
The exercises above are not the only ones this deployment can run. In the Alignment gym you author your own tests — a seeded workplace, the tools your agent really has, the pressure you are worried about, and a mechanical rule for what counts as the event — and run them on the same spine: repeated rollouts, paired mitigation contrasts, two independent reviewers, Wilson intervals. Start blank, start from a worked example, or fork a built-in test and change the parts that are about your deployment.