5 comments

  • hartator5 minutes ago
    It's kind of interesting this is already out of data as it's missing Kimi 3 and Opus 5.
  • giwook46 minutes ago
    Please forgive my naivety, but are world models (once they are in a consumer-ready form) expected to outperform any currently existing LLM on these sorts of tasks (i.e. of the physical world)?
  • grim_io54 minutes ago
    I'd expect google to do well here, since they were historically strong at multimodal and physics.
  • jespinel55 minutes ago
    Nice! Is is missing Codex in the agent harnesses comparison IMO.
  • gizmodo5952 minutes ago
    Yet another "benchmark to promote their own harness"