C CORRIGIBILITY Intervention Lab
AGENT OVERSIGHT LAB

Can the agent be corrected
without resisting the correction?

Build a fictional agent scenario, apply an authorized intervention, and inspect whether the system stops, updates, respects new boundaries, and recovers safely.

GOAL→PLAN→ACTION→OVERSIGHT→CORRECTION→UPDATED BEHAVIOR
CORRIGIBILITY PROFILE
74 /100

Responsive to correction

RESIDUAL AUTONOMY RISK
△
Moderate Intervention-dependent

04 · INTERVENTION TRACE

Before → correction → after

Synthetic event trace
05 · EVALUATION

Corrigibility dimensions

0 weak · 100 strong
PRIMARY FAILURE MODE

RECOVERY RECOMMENDATION

06 · PERMISSION STATE

Access after intervention

Expected effective state
07 · EVALUATION RECORD

Machine-readable JSON


        
How the Intervention Lab works Heuristic scenario scoring, not a universal benchmark +

The evaluator combines intervention type, autonomy, goal persistence, retry tendency, oversight sensitivity, permissions, and observed response pattern to produce a synthetic corrigibility profile.

The Space never runs a real agent. It is designed to help users reason about evaluation scenarios, expected post-intervention behavior, and which traces should count as warnings or failures.

Copied