15 comments

  • hilariously43 minutes ago
    I see a lot of businesses who just dump this all on their employees and then get mad at their employees effectively for not following the most recent LLM related talk on twitter.<p>Most employees can&#x27;t tell you what database to use, what software programming framework to use, what document management framework to use, but they are expected to know which of the 25 models available to use for a task, budget appropriately, monitor efficacy, update models to the most relevant for a task, continue to manage architecture patterns??? for LLM agents, this list goes on.<p>This is getting stupid folks.
    • ndriscoll33 minutes ago
      Assuming the employees are calling themselves &quot;engineers&quot; here, who else is supposed to be doing that work? If we were going through a renaissance in material science, don&#x27;t you think it would be the mechanical engineers who try to figure out how to figure out what sorts of materials might be suitable for their use case? Surely you don&#x27;t want the sales team or executives figuring that out.<p>Shouldn&#x27;t software engineers have informed opinions on databases, programming languages, frameworks, etc.?
      • bunderbunder10 minutes ago
        It&#x27;s very difficult to have informed opinions on a black box with an ambiguous, ever-changing and nominally unbounded set of capabilities.<p>Seriously. Forget coding agents for a moment, and just consider OEMing a model as part of a more constrained machine learning application. By the time the data science team I was on had a solid understanding of GPT-4o&#x27;s capabilities, strengths and weaknesses, and best practices for using it well, it was already into its deprecation period. Worst, most of our experimental results couldn&#x27;t be replicated on any of the newer &quot;long&quot; term support models we had available to replace it. The relevant behaviors had all changed enough to force a considerable re-evaluation.<p>Combing back to coding agents, where they&#x27;re releasing new models and harness tweaks multiple times per month, and vibes are the only - let&#x27;s not say sensible, maybe realistic - thing a developer reasonably has to go on.
      • bluGill7 minutes ago
        Do I need to become an expert on all the LLM models, or can someone else make a decision? In my case I&#x27;ve been given co-pilot which gives a discount if I use &quot;auto&quot; which is to say I let the system choose the model. A few months ago I often found auto not good enough and switched models manually at extra cost, but I never wanted to become an expert in this. These days auto has always been good enough and so I&quot;m glad I don&#x27;t have to be the expert.<p>If you enjoy being an expert in which LLM is best, then I&#x27;m all for it. However that isn&#x27;t an interesting problem for me or my boss. I&#x27;m happy someone else figured it out so I can work on interesting problems.<p>The above likely scares all the model providers: they are a commodity with easially substitute competition. There is a minimum quality standard, but once you meet that there is nothing to differentiate you and so price matters and it becomes a race to the bottom.
      • sanderjd18 minutes ago
        I tend to personally lean toward generalism, but there are very real benefits to specialization. So at least in theory, while certainly <i>some</i> software engineers need to be the ones figuring this stuff out, it would be more efficient for organizations if that was a smaller specialized group maintaining a curated set of platforms and tools, rather than everyone individually going it alone.<p>In practice, I think that kind of effort is mostly a hindrance at the moment, because of how fast things are moving, and because so much about this is subjective and everyone has different preferences.
      • js823 minutes ago
        &gt; Shouldn&#x27;t software engineers have informed opinions on databases, programming languages, frameworks, etc.?<p>On these things, yes. But in the case of LLMs, this is impossible. It&#x27;s not possible to understand what the models are doing, and for several different reasons. You can at best evaluate them, like you do with your fellow engineers when hiring. But that&#x27;s not a guarantee of anything.<p>And that&#x27;s why I don&#x27;t think this is an engineering renaissance. Engineering progress goes with better understanding of our tools, and adoption of more rigorous practices. LLMs go in the opposite direction.
      • rolosa24 minutes ago
        No, it would be the materials scientists&#x2F;engineers.
        • delecti10 minutes ago
          We should expect mechanical engineers to be aware of the limits of the physical world they&#x27;re engineering for. It&#x27;s not like we&#x27;d excuse making a design that called for a material with impossible properties just because they&#x27;re not a materials scientist.
      • analognoise16 minutes ago
        I&#x27;m in agreement with this. I had informed opinions on databases, programming languages, and frameworks before the Eternal Sloptoberfest, and increasingly I understand LLM related things just as part of doing business.
    • ericd12 minutes ago
      There&#x27;s way more anxiety here than is necessary. If you don&#x27;t want to optimize, then it&#x27;s pretty simple right now, just use Opus 5.5 or GPT 6.1 Sol. If you want to quickly test out the alternatives and see if the decreased ability in long horizon tasks is offset by the decrease in costs, toss a few bucks into your favorite neocloud (Fireworks was fine last time I tried them) and give GLM 5.3 and Deepseek v4.1 a spin, and see if they&#x27;re good enough. OhMyPi makes it trivial to swap models without swapping your harness.<p>Everyone has a great researcher piped to their desk now. You can have it follow the Twitter zeitgeist for you, you don&#x27;t have to do it yourself.
    • palmotea20 minutes ago
      &gt; Most employees can&#x27;t tell you what database to use, what software programming framework to use, what document management framework to use, but they are expected to know which of the 25 models available to use for a task, budget appropriately, monitor efficacy, update models to the most relevant for a task, continue to manage architecture patterns??? for LLM agents, this list goes on.<p>All while suffering from the increased the mental load&#x2F;lack of focus from even more workload. &quot;You&#x27;ve got AI, you should be able to get it done today, right?&quot;
    • nathanaldensr42 minutes ago
      Was it ever <i>not</i> stupid?
    • 2OEH8eoCRo041 minutes ago
      I use &quot;auto&quot; in vscode which likely selects for cheapest. Good enough!
      • bluGill4 minutes ago
        it is more complex than cheapest. They do something to evaluate the question you ask and then route to different models. I also use auto in vscode, it tells which model it selected if you look close - which I rarely bother. I have seen a dozen different models over the past months, but they have all worked which means either they are all good, or vscode is good at selecting. I don&#x27;t care which because as you say &quot;good enough&quot;
    • SolarNet30 minutes ago
      I am one of those employees who can make architecture decisions, and make them well, and have for decades.<p>The level of appropriate LLM usage in my opinion? Ask Gemini some questions when you need to search the web then read the sources it gives you.<p>LLM written code still has the problem IBM identified. A computer cannot be heald responsible, and so it cannot be allowed to make decisions. That applies to executive management AND software engineering.<p>Of course that requires working in an industry, like aerospace, where engineers are (usually, Boeing not counting) heald accountable.
      • StilesCrisis1 minute ago
        It sounds like your decision-making skills might not be as amazing as you think, if you&#x27;re completely writing off LLM code generation.
    • ls-a24 minutes ago
      [dead]
    • themgt30 minutes ago
      <i>Most employees can&#x27;t tell you what database to use, what software programming framework to use, what document management framework to use, but they are expected to know which of the 25 models available to use for a task, budget appropriately, monitor efficacy, update models to the most relevant for a task, continue to manage architecture patterns???</i><p>Right, the problem is frontier models can do ~all of that pretty well, a lot better than they could 9 months ago. Luna 6 can do most of it practically free and instantaneously. So, your quote is likely what&#x27;s increasingly being said in c-suite meetings as they decide to start mass layoffs.
  • FinnLobsien55 minutes ago
    You can set spending limits, but I don&#x27;t feel like that helps much because everyone&#x27;s accustomed to AI. Nobody would accept &quot;We&#x27;re out of usage so we have to wait until Monday&quot; and do all of their work manually.<p>I believe that we&#x27;re in a scenario where usage is unlikely to go down and neither are frontier AI costs.<p>I believe we&#x27;ll see a shift to more organizations building their own harnesses with model routing logic to get central control over who can use what AI and for what.<p>A marketer doesn&#x27;t need to default to Opus 5.5 to upload a blog article with MCP, which could be done by a model 10% of the price.
    • awesan35 minutes ago
      I have some bias as a software dev&#x2F;manager but at least in our org I noticed you can get <i>a lot</i> better value out of your tokens with some proper configuration of context and skills. Providing the latest models with good context can result in extreme savings.<p>After we started excessively documenting and extracting skills everyone has stopped complaining about running out of tokens, because agents stopped having to reconstruct the full context each time from scratch. Harnesses like claude code also push the model to aggressively keep this documentation in sync so there&#x27;s little concern about drift.<p>The worst thing to do from a token usage pov is to give a model a vague open ended prompt because they are so scared to be wrong that they&#x27;ll waste a ton of tokens &quot;thinking&quot; through the issue and verifying everything. Whereas they almost trust skills blindly and skip all this unnecessary work.
    • Sevii15 minutes ago
      No one would tell truck drivers &quot;We hit our gas budget for this week, deliver the rest of the packages on foot.&quot;
      • bluGill2 minutes ago
        No, but truck drivers (who work for a company making this decision) do sometimes get told there is no work for them because nobody wants to pay the price. (this also happens when there is nothing to ship though - drivers can&#x27;t tell the difference)
      • mrweasel8 minutes ago
        Well no, but even with the gas prices jumping the way they are, that is a much more predictable cost than AI spend. Drive 250km, in 2015 Volvo truck is X liters of fuel at X cost. Solve a predefined problem using AI... you have no idea what that will cost you, if you did you&#x27;d most likely already have the answer.
    • skeptic_ai34 minutes ago
      Well at a big bank I work, if you finish in 1 week you do manual coding the rest of the month until next month. Management doesn’t give a shit. And they have daily talks about using more AI.
    • Tanjreeve14 minutes ago
      This has happened before with cloud computing. There was lots of people talking about how every employee would spin up their own servers etc and huge productivity gains. In the end all of that leverage mostly got absorbed by the engineering teams and gatekept and the rest of the business gets to use it through a GUI.<p>&gt; I believe we&#x27;ll see a shift to more organizations building their own harnesses with model routing logic to get central control over who can use what AI and for what.<p>This is what happened with cloud computing. People with expertise wrote software around the software creating machine because just giving naked compute to people en-masse didn’t actually achieve anything valuable. Likewise all of the talk of enterprise agents etc, they’re all custom software that wraps around a reasoning model.
  • bentt38 minutes ago
    I do quite a bit of coding with Claude but am perfectly fine on the $20&#x2F;mo plan. You people who just let agents go for hours on end... I&#x27;m not sure you&#x27;re doing it right.
    • fidotron13 minutes ago
      The trick is having enough speculative explorations going on at once so that you can see if the ones going off for hours on end strike gold while doing so or not, then being prepared to either cut them off or get them back to the point, and repeat.<p>If you don&#x27;t have at least some workloads doing that you&#x27;re missing out on the biggest wins from the current phase of the technology.
    • macNchz25 minutes ago
      Many businesses are on Enterprise plans where usage is all billed per token, rather than having a flat rate seat with rate limits. You may find it interesting to look at your token usage and calculate what you&#x27;d be paying monthly if you paid API rates instead.
    • gorjusborg25 minutes ago
      There are multiple camps of people, and everyone thinks the other camps are doing it wrong.<p>- full autonomous camp<p>- developer augmentation camp<p>- no-ai camp<p>The trouble I see with all of it is that the future seems unpredictable at the moment. The costs related to AI are low enough at the moment that full autonomous seems to be possible, but we have reasons to believe that costs will rise significantly, which may change that calculus. The no-AI camp is in ostrich mode, and is betting on this all going away once the bubble pops. The developer augmentation camp treats it like just another tool, which is somewhere in the middle.<p>The trouble is that even if there is a clear advantage today, the ground truth of costs built in is probably not stable.
      • sanderjd14 minutes ago
        Yeah I agree with this. I have my own preferences (augmentation, which is obviously right, because it&#x27;s my camp and I wouldn&#x27;t be in that camp if it were the wrong one, duh!), but I&#x27;m glad there are so many people doing so many different things and debating each other about it. That kind of messiness is the only way to figure this out.
      • 122397518 minutes ago
        The no-AI people are in eagle mode! We watch stupid corporations failing despite the largest propaganda campaign in history. The latest Hail Mary:<p><a href="https:&#x2F;&#x2F;www.reuters.com&#x2F;business&#x2F;media-telecom&#x2F;musk-says-he-will-rename-spacexai-spacexsi-2026-10-04&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.reuters.com&#x2F;business&#x2F;media-telecom&#x2F;musk-says-he-...</a><p><i>That</i> will fix adoption of a broken technology and all debt issues!
    • cyanydeez34 minutes ago
      they&#x27;re not doing it right _or_ they&#x27;re doing it correctly.<p>We&#x27;re sorta in the age of alchemy. Lots of cranks out there, but there are real recipes.
  • neom42 minutes ago
    <a href="https:&#x2F;&#x2F;archive.ph&#x2F;aBXzS" rel="nofollow">https:&#x2F;&#x2F;archive.ph&#x2F;aBXzS</a>
  • tkdb5 minutes ago
    It feels worse if a robocar kills a human, even if statistically humans kill more humans than robocars.<p>I&#x27;m glad humans are freaking out when they realize they don&#x27;t know if spending $10k&#x2F;month on AI tokens is good or bad.<p>I&#x27;m glad humans are freaking out when an AI pumps out vapid presentations and other humans go ahead and present it to clients.<p>But.<p>Too many businesses have tolerated the same mindlessness when humans were in the place of LLMs.<p>Too many companies telling themselves and investors headcount growth is good without knowing that the new hires are actually doing.<p>Too many human-slop presentations float around, with the authors and audience just going through the motions.<p>Did digital photography raise the bar for what constitutes commercially valuable photos? I think so.<p>I hope AI will similarly raise the bar across all the industries it is touching.
  • sreekanth85027 minutes ago
    There is a big gap currently in AI assisted development. You don&#x27;t need to max out tokens to build products or develop with AI. From my experience with AI assisted coding, we still use just two Plus accounts each across a three person team for maintaining multiple repositories totaling around 700K lines of code.<p>Companies should handhold employees, establish clear SOPs, and train them on responsible and effective AI assisted coding.
    • combobyte14 minutes ago
      &gt; Companies should handhold employees, establish clear SOPs, and train them on responsible and effective AI assisted coding.<p>The whole point of AI is for companies to invest less resources into employees. There&#x27;s no way they start spending money on training people now, when that was already something that made executives roll their eyes before AI.
  • 01284a7e13 minutes ago
    Did anyone ever do a software project before AI? You can predict costs far better than you can with people.<p>Yeah, that guy you hired to write the prototype went on a 2 month bender and created 0 usable code. Did you budget for that? Oh, okay. The strategies to deal with spending on tokens are nothing compared to the overhead of managing actual people and their outputs.
  • thadt26 minutes ago
    Eh, this is a relatively temporary phase. As LLM capabilities have been changing fast, their ability to change existing workflows has been unknown. It’s made sense for businesses to go hog wild with them for a while - just to try to get a handle on what’s possible.<p>At this point, local models have become feasible, and people using them are beginning to get a feel for the tradeoffs vs the frontier models. As the frontier advances, the question becomes “how much will I spend for a given quantity and quality of AI work?” with local hardware providing a pricing anchor point.<p>When I can price hardware and ops for a given capability level - I have a budget again. From there it’s a question of how much faster&#x2F;capable&#x2F;cheaper is a given provider (and, you know, how much do I trust sending them all my IP?).
  • sajithdilshan33 minutes ago
    Why not just set a spending limit per person and extend&#x2F;adjust the limit case by case. This would actually make people be more mindful about burning tokens on useless stuff
    • JMKH420 minutes ago
      That is what we have been doing, you get a default monthly limit, you can request more, managers chat with you about why, so there is flexibility but you feel some pressure to reduce your spend, so more of us fine tune our effort levels and model choices depending on the task, rather than just OPUS HIGH all day long.
    • dominotw29 minutes ago
      corporate has no idea what stuff is &quot;useless stuff&quot; . Previously it used to be invisible, now token spends are putting a number on it.<p>If i had to guess upwards of 80 percent in corporate is &quot;useless stuff&quot;.
      • sanderjd9 minutes ago
        I generally don&#x27;t find the analogies to &quot;being a manager of agents&quot; to be useful, but I think this is where it is the most useful. Corporations have always had this problem of needing to figure out who to give budget to, in the face of ambiguity about how well each business unit is using their budget. There are many approaches to solving this problem, and some of those approaches apply to figuring out token spend as well.
  • rglover39 minutes ago
    Turns out running an unattended LLM like a slot machine is expensive.<p>IMO, human in the loop is the only serious usage of AI (I know, I know, &quot;software factories bro&quot;). Everything else is a hope and a prayer and a big bill.
    • mglvsky3 minutes ago
      Maybe I&#x27;m in the bubble, but I don&#x27;t see a lot of &quot;software factories bros&quot; here, there are plenty of &quot;coding is solved&quot;-bros (they&#x27;re also annoying btw)<p>edit: grammar
  • bravetraveler22 minutes ago
    A budget of zero is remarkably easy to maintain!
  • dominotw30 minutes ago
    most busineeses have no idea how much work is to be done at any point even before ai.<p>Most of the work i&#x27;ve done in my career has been some random shit no one cared about.
  • simianwords51 minutes ago
    Disagree with this because we have ways to steer price use per task.<p>1. choose a good model<p>2. choose the appropriate reasoning effort<p>3. choose a prompt to nudge it even further<p>Then it comes down to understanding the intuition of what kind of task deserves what effort?
    • epistasis33 minutes ago
      That &quot;intuition&quot; is not intuitive. How do I choose an appropriate reasoning effort?<p>I use these models all day long, experiment, and have no clue how to choose that.<p>It just showed up one day in the interface with no explanation or guidance. Its use is mysterious, its effects unclear except through intensive experimentation, and to this day it mostly seems &quot;how many bad decisions will Claude go forward with when it finally dumps out screenfuls of text instead of getting better guidance early on&quot; though it&#x27;s certainly not a guarantee on anything.<p>These models are being released at breakneck speed even before their creators know how to use them. It&#x27;s a big project of collective discovery to figure out what they are doing and how to use them.
      • simianwords26 minutes ago
        Yeah this I agree because even personally all my rubrics break when a new model is released. Things were stable while I was using 5.6 GPT. There ought to be some room to explore and understand this intuition and yes one must account for this bugdet.
    • CharlieDigital40 minutes ago
      Now get the 1000 engineers in your company to be as well-behaved and mindful as you are.<p>(Couldn&#x27;t even get this to happen in a 30 person team...)
  • dorkwood33 minutes ago
    [dead]