Evaluation is an always-running service, not a testing phase.
A
B
C
D
Online Telemetry detects drift/errors
Triggers HITL review for edge cases
Human feedback feeds into Offline Datasets
Offline Scenario Evals validate system updates before pushing back to Production