terminal-bench: benchmarks for ai agents in terminal environments
terminal-bench is a collection of harbor-native benchmarks to help agent makers quantify their agents' terminal mastery
terminal-bench 3.0 is now in development
terminal-bench-science is now in development
I want to test my agent
AI text/layout recreation from video frame; verify against source image.