A challenge with this kind of study is that coding agents (Claude Code, OpenAI Codex) only started working really well in late November, which for most people meant early January due to the December break.<p>General agents (OpenClaw, Anthropic Copilot, ChatGPT "Work") started working even later than that.<p>This category of software may have a much more meaningful impact on work than the mostly-chat systems we were using from 2022-2025.<p>Studies that mainly focus on 2022 to end of 2025 might be missing out on a material uptick in capabilities.
Anecdotally, I (along with the rest of my team) got laid off from a big tech company at the end of January. The stated reason was "AI", although I think as we all know the real reason was "we overhired in 2022". I managed to get a job by April without putting in a lot of work, though (mostly just listened to recruiters until something hit my fancy). My work experience is good, but I don't know if it's so good that I would be immune to market effects. The company that hired me is still hiring other software developers, and I was told that it took them a long time to find someone with my qualifications. (We use Claude, although so far I haven't seen a lot of AI psychosis like I have in other places.. that was also a big filter in my job search). Long story short, I think narrative of job loss is way overblown. It's more doom trolling from anthropic and openAI and their mouth pieces as far as I can tell.
It's so early game, is the real problem. My experience so far, has been that it's barely a junior dev. I've met so many in my career that think reading stackoverflow, or watching a youtube video mysteriously makes them an expert.<p>Well anyone can use (prior to AI) a simple linter and learning to code isn't that big a deal. It's learning the pitfalls, the traps, that's the issue. And so far Opus just seems to fall into them again and again. I guess the best way to put it, is that it's not an architect. I sees no big picture, and that's not really a surprise with (compared to a human) an incredibly small context window. When I'm on a project, or working with a codebase, I often have <i>years</i> of "context window". And I have a career of "don't do this" context window.<p>So what I wonder is, will this be resolved? Will that awareness of larger scope be solved? If <i>that</i> happens, we'll be in another ballpark of competency.<p>Some companies have massive codebases. Are these companies slowly gaining rot in those codebases, a swiss cheese effect, which eventually will result in collapse? Because I've worked where a bad hire had just this effect over time. And what I worry about isn't using Claude to speed one up, it's the DEV that uses Claude and just "meh" and submits because it passes regression + other tests, and then a manager or code reviewer uses Claude and "meh" because it's a pass too.
> A challenge with this kind of study is that coding agents (Claude Code, OpenAI Codex) only started working really well in<p>2026? 5? 4? 3?<p>Heard this one way too many times.
I guess it's important who one hears this from.<p>I just spoke to a fried who is a headhunter and who's been trying to automate his processes for a while (he likes to fiddle and certainly has skills, but he's not an engineer). He kept trying, but it just wasn't good enough.<p>Now he said with GPT Work and Sol, it worked, but the key point is: all of it suddenly worked.<p>The problem was one of reliability, of handling edge cases. All previous attempts / model-harness-combinations were too brittle and needed too much observation and fiddling - cheaper to do it yourself.<p>Now he says "I don't know why I would ever hire a recruiter [the folks doing the cold outreach] again. I can focus on the candidate screening and acquiring projects, everything else is fully automated".<p>This doesn't come from an engineer or an AI lab, but a technically inclined power user, and I think this is where things get interesting.
Then it just becomes a new baseline (everyone have access to the same LLMs), and recruiting moves up the philosophical ladder where human can add more value. What will it be? I don't know, I'm not a recruiter.
I get the point having read much the same from Tesla (and fans) regarding self driving cars that still haven't done half the things that Musk said was just around the corner pending regulators a decade ago and repeatedly since then.<p>And myself I keep making comparisons between AI and the progress in 90s video games where every minor improvement got called "photo realistic" and then forgotten with the next game engine: <a href="https://archive.org/details/nextgen-issue-26" rel="nofollow">https://archive.org/details/nextgen-issue-26</a><p>So I'm not gonna say "this is it" when the software quality really matters, and I absolutely won't speak to progress (or lack of it) outside of software.<p>But I will say "you can look around and easily see small businesses using AI to generate posters, quite a lot of small business software and websites are in the same category: the mistakes are real but increasingly don't matter".
>the mistakes are real but increasingly don't matter...<p>I think it would start to matter once again. People will get fed up of AI posters and art. I think they already are...and once some threshold is crossed, the business won't dare to use AI generated assets/designs.<p>Turns out humans are much better at recognizing patterns in stuff that is generated ONLY using patterns from human generated content.
> People will get fed up of AI posters and art. I think they already are<p>Agreed, but will this look like a meme/fashion cycle? If so, re-prompt each year with a different look. Yes, there are still issues here, a friend found an image he was amazed was AI generated, but to me it was obviously so, so I showed him a screenshot of ChatGPT making something just it and included my prompt:<p><pre><code> create image: hand drawing of cute springer spaniel puppy looking sideways, various geometric shapes drawn in layer behind and in front of the puppy, all done in style of 7 year old using crayons with mediocre colouring-in skills
</code></pre>
As I said to them:<p><pre><code> yeah, the line thickness feels AI, to me, the bad colouring-in scribbles feel like just the art style it was propmpted with
it's like: it gets the big picture of the composition, and it knows how to colour in badly, but it doesn't know how to draw a dog as badly as the colouring in
</code></pre>
> Turns out humans are much better at recognizing patterns in stuff that is generated ONLY using patterns from human generated content.<p>We're better at recognising patterns full stop. All biological brains are, and needed to be better than the current state of the art in machine learning because if a living organism was as poor at learning patterns as the SotA in machine learning, the organism would starve to death before being able to pick up anything and eat it.<p>AI also has a second disadvantage, because there are so few models: the laziest of ChatGPT "thinkpiece" blog posts being everywhere is hard to miss, and 5000 fake bloggers all prompting the same model with "find biggest news story of today and write a blog post about it in a way that maximises my ad revenue" will get 5000 almost identical posts. This will remain true while each instance of the most commonly used AI fail to talk to each other in a way that at least mimics them collectively getting bored with writing the same thing 5000 times, it does not depend on e.g. quality.
It seems to be true this time though; I have observed it myself and heard it from several experienced developers I personally know and respect. It feels like some threshold was crossed with Opus 4.5 and Gpt 5.3, where the models are now able to reliably solve certain classes of problems that were previously unreliable.<p>Time will tell of course, and it’s early, but inflection points do exist with progress.
Perhaps the LLM companies need to start hiring true Scotsmen?
Nobody was saying coding agents started working in 2023 or 2024, because the category was defined by Claude Code which was first released in February 2025.
Claude 4.5 was it (nov 2025?), without a doubt. It went from frequent hallucinations to highly usable with much less garbage output. If you were making demos of AI tools around this time your demo/pitch/product was saved and you probably looked like a genius.
Yep, the goalposts just keep shifting. In reality: they still don't work well, unless you're content with producing low quality work.
This is the opposite of my experience since about February of this year.
The quality of the output is so variable. It depends on the model, “effort level”, prompting, probably even the programming language/app functionality, and libraries involved. For example, I find LLMs are best at making simple web apps. These web apps, while simple, would still take a senior engineer perhaps a week or two to create, but LLMs can spit them out inside of an hour. Conversely, LLMs struggle with things like Docker or local model stuff. Parallelization of code is a mixed bag. In these areas I think it often would have been faster for me to write the thing by hand.
"unless you're content with producing low quality work." - With the right guiding hand, it is a productivity multiplier without compromising quality. As a fully autonomous developer, it is a disaster.
> With the right guiding hand, it is a productivity multiplier without compromising quality<p>This just reads like another variation of “it’s the user not the tool,” which is just endless runway for always blaming people and never acknowledging the limitations of LLM’s.<p>I’d be curious to hear how the recipients of your work enabled by the “productivity multiplier” feel about the quality.
[flagged]
I produce code that is significantly higher quality with the assistance of coding agents, because I no longer succumb to the temptation to cut corners due to lack of time.<p>One example: everything I do is properly tested and documented now, even the most trivial of changes. Previously I would have weighed those tradeoffs and sometimes decided not to bother with the tests because they weren't worth the time.
What kind of work are you doing and what do you consider to be quality or not?<p>Of course don’t let me assume, maybe you have a higher quality disproof for the Jacobian conjecture you could share with the class.
Anecdotally I'm seeing a lot more recruiter activity/interest now than this time last year.<p>But it seems more correlated with hype-cycle-stage than anything else. Right now a lot of founders seem to be convincing a lot of VCs that they can make $LOTS by replacing/changing $BIG_INDUSTRY/$BIG_PRODUCT with an agent-first blah blah replacement, and then using that money to hire more people to manage/execute/coordinate the coding agents...<p>Last year, by comparison, there seemed to be a mood of "software will stay the same but will require less people" while right now there's a lot of hype around "we can build different types of software or build it in different ways" and those early-stage things are in growth-mode. That guarantees nothing about how many people they'd need in the future, or their success at all, ofc.<p>The news that I'm getting from contacts in non-startup-land is a bit different - still layoff threats. Still pressure to use AI tools more. Mixed confidence on whether or not longer-running "agent" modes are that much more effective-without-breaking-things in legacy code if not used with care.
However, companies have been using AI as an excuse for layoffs since well <i>before</i> January 2026, which corroborates the study's conclusion. (Source: <a href="https://layoffs.fyi/ai-layoffs/" rel="nofollow">https://layoffs.fyi/ai-layoffs/</a>) There is certainly an uptick starting 2026, but that could be explained either by AI <i>actually causing</i> more layoffs, or by AI <i>becoming an even better excuse for</i> layoffs.
This isn’t the first study showing this though. It’s pretty simple, programmers and IT were severely overhired during the pandemic, there are massive job losses now, and it’s easy to blame AI when in reality there are a lot of economic factors and AI isn’t increasing productivity as much as anyone would think.<p>Maybe the future will change that for very specific things, but I think people should be learning and preparing for that, which isn’t any different than what everyone has been told in every job market since the start of the Industrial Revolution.
I personally hope that AI continues not to result in a noticeable negative impact on employment and that pandemic over-hiring turns out to be the major factor for all of the layoffs.<p>I'm nervous that the studies which show that so far don't seem to be taking the 2026 improvements in coding and general agents into account.
>programmers and IT were severely overhired during the pandemic<p>1) I hesitate to believe that losses were disproportionately technical roles as opposed to administrative.<p>2) Over-hired by what metric? It's well known that hiring never fully recovered after the GFC; was the recruitment post-pandemic just bringing us to parity with where we had been 20 years earlier?<p>Not to say that I disagree with your following point. The AI overspending and the layoff cost-cutting are not in a direct causal relationship; both are rather symptoms of a common corporate pathology.
This. Also I'm finding out that, after being blown away by agent mode lately, non-agent mode still kind of sucks across frontier models. Using GPT and Gemini in non-agent mode is asking for inaccurate information confidently presented as the truth. Turning on agent mode fixed a lot of that for me.
For whatever anecdotal evidence it's worth - that has been my experience as well. For the first time last December I noticed the harnesses performing like actual workers.<p>Everything impressive has happened in the last six months.
It does feel like we are in a transition period, and its not clear what conclusions we can draw about any sort of "steady state" yet.
Any research on the impact of AI would have been lagging indicators. It's not the researchers' fault[0], but the nature of a field moving at neck breaking speed. Remember there was a paper saying programmers were 20% slower with AI?<p>[0]: well...
That early 2025 METR study was particularly interesting because participants self-evaluated themselves as 20% faster, but the measurements showed they were actually 19% slower.<p>All the reports of productivity since then are self-reported, or using questionable measures such as SLOC and PRs, so it’s reasonable to say that productivity improvements are still unknown.<p>Unfortunately, METR hasn’t been able to replicate the study because they couldn’t find enough willing participants.
Well… what
And this is why the labs cannot just "stop training and become profitable". I can't imagine they would like it if studies like this will actually become credible.<p>"Move fast like a blur so people can't see that you have no clothes"
Well, in addition to that you also have to consider that companies are now paying (increased) API pricing. It would be an understatement to say that my clients in the F100 range are skeptical at best regarding the gains they’ve seen compared to the costs.<p>This has led to many of them instilling dollar limits or demanding proof of increased productivity (not just output) with the implication being if you don’t provide value with it it’s getting taken away.<p>So that is to say, if they aren’t happy with the price now, how will they feel when it goes up again compared to just keeping a certain headcount?
<p><pre><code> > So that is to say, if they aren’t happy with the price now, how will they feel when it goes up again compared to just keeping a certain headcount?
</code></pre>
that got me thinking: how are companies expensing ai costs? as personnel expenses or r&d etc?
>This has led to many of them instilling dollar limits or demanding proof of increased productivity (not just output)<p>They should have done that <i>from the beginning</i> - demanding proof of increased productivity - <i>if</i> that was their goal. otherwise they were not using their brains well enough.<p>And you <i>doubly</i> don't want to work with them, first because they confused output with productivity at first. and second, because they're parroting the productivity metric.<p>You only need one guess for whose pockets the productivity benefits go into.<p>10 . 9 . 8 . 7 . 6 ...