6 comments

  • smithclay8 minutes ago
    More benchmarks comparing effectiveness of various AI SREs is welcome and overdue: especially vendor-vs-vendor comparisons. One area that I think is going to be really interesting and important is the best way to emulate complex IT environments for evals.<p>Some related work I recommend checking out: - <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2609.33023" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2609.33023</a> (new last week!) - <a href="https:&#x2F;&#x2F;github.com&#x2F;SREGym&#x2F;SREGym" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;SREGym&#x2F;SREGym</a> - <a href="https:&#x2F;&#x2F;github.com&#x2F;hyperdxio&#x2F;hyperdx&#x2F;tree&#x2F;main&#x2F;packages&#x2F;hdx-eval" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;hyperdxio&#x2F;hyperdx&#x2F;tree&#x2F;main&#x2F;packages&#x2F;hdx-...</a> (Clickstack&#x27;s version) - <a href="https:&#x2F;&#x2F;github.com&#x2F;grafana&#x2F;o11y-bench" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;grafana&#x2F;o11y-bench</a> (Grafana&#x27;s version)
  • tkkiran16 minutes ago
    Cool, as I see your product react proactively rather than Claude being reactive so that with your tools automated investigations happens without me asking explicitly the problem and root cause as I understand? Do you guys also have mitigations?
    • emrahsamdan10 minutes ago
      Yes depending on which connector you chose to connect. There could be a new PR on GitHub, a rollback in CircleCI or a feature flag switch on LaunchDarkly. All waits for a human to chime in by default.<p>These are all becoming tablestakes, imo for any AI SRE. The real moat is to build the data layer in an efficient manner so that the investigation (and mitigation) is fast and cost-efficient. The models are making it easier for any company at the same time.
  • nikhilunni1 hour ago
    Cool that you guys created this, but interpreting the results: why wouldn&#x27;t I just use Claude instead of an AI SRE tool?<p>Or if there are some features of an AI SRE tool that make it better than Claude + some MCPs, should those be captured in this same benchmark?
  • rafiyashaheen111 hour ago
    How much time did it take to build this? Loved it! What problem does this solve?
    • emrahsamdan57 minutes ago
      Running the test takes around 5-6 hours per vendor depending on how much wrestle it requires (not every system is ideal for ai agents - including ourselves) but the actual work was to come up with the real world scenarios and what would be the expected RCA and remediation for this. Our own SRE team worked for 1.5 weeks for that.
  • esafak47 minutes ago
    Emrah, it seems that EDX does not offer any edge over GCX, when paired with an agent like Claude?
  • tanmoy113933 minutes ago
    [flagged]