4. Agentbench
tests a model's ability to reason, plan, and use tools over extended, multi-step workflows, rather than just measuring static question-answering accuracy
AgentRTX 61.3
Mem0 58
RAG 47
Baseline 35.7
AgentRTX 98
DSPy 82
Baseline 57
AI text/layout recreation from video frame; verify against source image.