AWS launches aws-bench: open benchmark for AI agents
🔍 AWS today announced a research preview of aws-bench, an open-source benchmark designed to measure how accurately and efficiently AI agents complete real-world AWS tasks. The suite includes test cases derived from actual AWS usage—such as investigation, troubleshooting, and infrastructure creation—pairing natural-language queries with defined resource states and ground-truth answers. A CLI tool is included to instantiate test environments, run evaluations, score results, and reset state, and the project is available on GitHub.
