4 comments
Terminal Bench 2/2.1 appears to be almost solved, so we probably shouldn't look too hard on that? Their Terminal Bench 4 score is soso.<p>Still quite impressive though.
Related:<p><i>Cognition launches new SWE-2 model</i><p><a href="https://news.ycombinator.com/item?id=49645443">https://news.ycombinator.com/item?id=49645443</a>
Not really news, terminal bench 4 is the new metric. It's only a few points ahead in terminal bench 4 of some open weight models you can run on a 256GB system.
These tests are pointless when often models get nerfed few days after release. Astra today is way dumber than just few days ago.
Is that something we have credible evidence for? Do they serve a better model when artificialanalysis (the benchmark site) is making the requests, and so on?
It is my own anectodal. Few days something I worked on usually got one shotted or got quality result. Today it is very much going nowhere and is stuck in reasoning loops.<p>Probably someone should build nerf tracker, because this is quite common that models get substantially worse once PR hype wears off and they quantise them more or simply route requests to older models with system prompt changed to say it is Astra and not Sol etc.
Not Astra but Sol hasn't been nerfed: <a href="https://marginlab.ai/trackers/codex/" rel="nofollow">https://marginlab.ai/trackers/codex/</a>
How and why do they get nerfed? To save money?