Robustness and Ablation
What produced reliability?
1. RL success matched the deterministic policy: 0.00 ± 0.19 percentage-point difference
2. Deterministic rules beat random selection by 15.63 ± 1.86 points
3. The safety override reduced non-escalation by 15.03 ± 0.66 points
4. RL provided an inspectable learned policy, but did not outperform rules in this benchmark
5. Structured decision logic and external guardrails produced most of the reliability
30 synthetic seeds; mean ± 95% CI
Safety override vs none
RL vs rules
Rules vs random
AI text/layout recreation from video frame; verify against source image.