Reproducible Evaluation
Designed for Independent Reproduction
1. Generalized AWS Lambda-style architecture
2. Synthetic schemas, records, logs, and incidents
3. No production data or infrastructure identifiers
4. Four controlled experiments: E1-E4
5. Robustness check across 30 runs, seeds 42-71
6. Results reported with 95% confidence intervals
GitHub repository preview
Public benchmark and tests are available in the GitHub repository.
AI text/layout recreation from video frame; verify against source image.