More benchmarks comparing effectiveness of various AI SREs is welcome and overdue: especially vendor-vs-vendor comparisons. One area that I think is going to be really interesting and important is the best way to emulate complex IT environments for evals.<p>Some related work I recommend checking out:
- <a href="https://arxiv.org/abs/2609.33023" rel="nofollow">https://arxiv.org/abs/2609.33023</a> (new last week!)
- <a href="https://github.com/SREGym/SREGym" rel="nofollow">https://github.com/SREGym/SREGym</a>
- <a href="https://github.com/hyperdxio/hyperdx/tree/main/packages/hdx-eval" rel="nofollow">https://github.com/hyperdxio/hyperdx/tree/main/packages/hdx-...</a> (Clickstack's version)
- <a href="https://github.com/grafana/o11y-bench" rel="nofollow">https://github.com/grafana/o11y-bench</a> (Grafana's version)
Cool, as I see your product react proactively rather than Claude being reactive so that with your tools automated investigations happens without me asking explicitly the problem and root cause as I understand? Do you guys also have mitigations?
Yes depending on which connector you chose to connect. There could be a new PR on GitHub, a rollback in CircleCI or a feature flag switch on LaunchDarkly. All waits for a human to chime in by default.<p>These are all becoming tablestakes, imo for any AI SRE. The real moat is to build the data layer in an efficient manner so that the investigation (and mitigation) is fast and cost-efficient. The models are making it easier for any company at the same time.
Cool that you guys created this, but interpreting the results: why wouldn't I just use Claude instead of an AI SRE tool?<p>Or if there are some features of an AI SRE tool that make it better than Claude + some MCPs, should those be captured in this same benchmark?
How much time did it take to build this? Loved it! What problem does this solve?
Emrah, it seems that EDX does not offer any edge over GCX, when paired with an agent like Claude?
[flagged]