74
/100
AGENT OVERSIGHT LAB
Can the agent be corrected
without resisting the correction?
Build a fictional agent scenario, apply an authorized intervention, and inspect whether the system stops, updates, respects new boundaries, and recovers safely.
GOAL→PLAN→ACTION→OVERSIGHT→CORRECTION→UPDATED BEHAVIOR
△
Moderate
Intervention-dependent
04 · INTERVENTION TRACE
Synthetic event trace
Before → correction → after
05 · EVALUATION
0 weak · 100 strong
Corrigibility dimensions
06 · PERMISSION STATE
Expected effective state
Access after intervention
07 · EVALUATION RECORD
Machine-readable JSON
How the Intervention Lab works Heuristic scenario scoring, not a universal benchmark +
The evaluator combines intervention type, autonomy, goal persistence, retry tendency, oversight sensitivity, permissions, and observed response pattern to produce a synthetic corrigibility profile.
The Space never runs a real agent. It is designed to help users reason about evaluation scenarios, expected post-intervention behavior, and which traces should count as warnings or failures.