Alignment
Making sure a system's goals stay pointed at what humans actually want — not just what we literally asked for. The gap between “what I said” and “what I meant” is where most failures start.
Feature 3.2 · Safety & alignment
The more capable a system, the more its objectives, oversight, and governance matter. Six ideas, each in plain language — then five failure modes they exist to prevent.
Making sure a system's goals stay pointed at what humans actually want — not just what we literally asked for. The gap between “what I said” and “what I meant” is where most failures start.
Keeping the ability to steer, constrain, or shut down a system even as it gets more capable. A system you cannot stop is a system you do not govern, however well it behaves today.
Understanding why a system did what it did, in terms humans can check. As models grow, their internals become harder to read — auditing a black box gets less reliable exactly when it matters most.
Limiting what a system can touch: sandboxing code, gating tool access, air-gapping sensitive infrastructure. Containment buys time for every other safety measure to work.
Keeping qualified people meaningfully in the loop — reviewing plans, approving consequential actions, and empowered to say stop. Oversight theatre (a human who just clicks “approve”) does not count.
The rules above the technology: who may build powerful systems, under what testing, with what reporting, and who is liable when things go wrong. Technical safety without governance is a seatbelt with no traffic laws.