A continual learning benchmark: Meridian Support Agent
A reproducible test-bed for continual learning in a tool-using support agent.
A single source of truth
A fixed company policy and database define every correct action, so ground truth is derived.
Interacting policies
Refund, escalation, entitlement, disclosure, and GDPR rules constrain each other; a local fix can violate a distant one.
Deterministic evaluators
code checks over the final answer and the tool calls
Decisions are tool calls
the agent's real action is which tools it invokes, not just its text
Regression-sensitive by design
tasks are arranged so over-fitting the latest fix is observable as a drop else
AI text/layout recreation from video frame; verify against source image.