WHY IT'S BROKEN
Two monitors that miss it
Behavioral testing
you'd have to guess the trigger
The model is correct on virtually every input you try
To catch it you'd need the exact trigger in advance — you won't have it
Joint cross-model features
the popular interpretability fix
Crosscoders learn features over base + fine-tuned together
The backdoor competes with everything the model represents — and gets buried
So where IS the signal? In what the fine-tuning actually changed.
AI text/layout recreation from video frame; verify against source image.