A lot of people here are responding to the message but not to the meaning.<p>It would be a good idea for young people to deeply know how these programs work. Not so that they can spend their career building them, but so that they can approach the next class of problems we'll all start trying to solve, with intuition all the way down to the weights and underlying mathematics. And also, to develop a healthy intuition of when "Just LLM it" will not be the right choice.<p>"Build an OS" wasn't a common university project because we were all expected to go out and work on Windows, but because understanding the bare-metal firmware for a computer helps you deeply understand how to intuit building for a whole class of problems.
Agreed. During comp sci we got to re-implement various algos of networks, OS, database, firmware.. and it all gave complimentary intuitions that were useful when tackling practical implementations and bottlenecks.
I'm not sure it's possible to have intuition about systems that work in thousands of orthogonal dimensions. In fact I'm pretty sure most of the research is people trying fairly arbitrary things and testing them and then post rationalising implied understanding of what is really happening on top of good outcomes.
I think it’s reasonable to have a shallow understanding of most parts and a deep understanding of a small number of parts. That’s how most engineers are.<p>Most software engineers do not have a deep understanding of CPU architectures. In fact they probably don’t even have a shallow understanding and get around just fine. How many of them are looking up the instruction set for the CPUs they deploy their CRUD app to in EC2?
But in the case of CPU architecture there are SOME people who understand how things work 100%, and they've built and vetted abstractions/mental models that enable other engineers and scientists to have that kind of mixed shallow/deep understanding in a way that works. On the side of LLMs we're still lacking an expertise which could flawlessly explain how these things operate; the abstractions that we're using are instead derived inductively and are totally unvetted.
You are right that the field doesn’t have a theoretically sound explanation for the architectural choices aside from “A works better than B”. However, I would argue this is an ideal opportunity for the “gentleman scientist” or eager 17 year old.<p>Basically every part of the original transformer was replaced with something more efficient or better:<p>LayerNorm -> RMSNorm<p>Sinusoidal position encoding -> RoPE<p>MHA -> GQA<p>ReLU -> GELU<p>What this means is that there is ample opportunity to improve on what we’ve done thus far.
In fact, one of the jobs of an engineer is to make sure that other engineers who don't work in his or her area do <i>not</i> need to understand that area deeply, yet build something reliable with it. They need just the summary that he or she writes up into the datasheet for the part. Ensure these conditions are met for safe/reliable operation, give it these inputs, expect these outputs, these timings, this energy consumption, this heat generation, frequency response, tensile strength, whatever.
I disagree. It's absolutely possible to develop an intuition about extremely complex mathematical ideas, including llms or high dimensional systems. Learning to build an llm is a great way to start building that intuition.
Developing an intuition about high dimensional systems is pretty different from understanding the character of some specific point on a 1e9+ dimensional manifold of parameters, in my professional opinion (setting aside all the degrees of freedom that come from the structure of the thing). Sure one can understand generic principles like the curse of dimensionality, but truly groking how an LLM works is basically an open problem as far as I'm aware. I'm not saying there's no benefit for amateurs to study how LLMs work, but let's be realistic about how far mere intuition can truly take <i>anyone</i> in this space.
What provable conclusions have you intuited around how LLMs work? Give me some examples to prove your point? I'm absolutely happy to change my view with enough data.
I'm not looking to change your view. But for other readers who are curious, here is a link to an interesting task to gain intuition. Ahmad is a good data point for someone who tinkered, built intuition, then started his own ai company.
<a href="https://twitter.com/TheAhmadOsman/status/2087742080793620593?s=20" rel="nofollow">https://twitter.com/TheAhmadOsman/status/2087742080793620593...</a>
I think we are talking about different things to be honest. Understanding LLMs is a fine thing to do but to say you understand what the network is actually doing, even in toy examples like MNIST is extremely tricky. Figuring out what a ~3 trillion parameter network is doing in thousands of dimensions seems intractable to me.<p>Linking to a Twitter thread about training a small LLM to be good at Wordle is not an example of what I'm talking about. It might well be a useful task but it doesn't allow us to understand deeply what's happening.
It is sometimes the opposite - a large number of things makes the system easier to predict and reason about (statistics, behavior of gases etc).
I'd have to push back, though not on the part you'd expect. Your description of human researchers is roughly right: a lot of the field is try-things-and-narrativize-after.<p>But the load-bearing assumption is that intuition has to be human-shaped intuition. Humans can't intuit thousands of orthogonal directions because we project everything down into a 3D metaphor and hope it holds. That's a fact about our hardware, not about the systems.<p>And the reason why is the most interesting part: nothing requires the compression step. A model or an agent can operate over the actual objects, holding thousands of runs and ablations in context and noticing regularities in the native dimensionality, without translating them into a picture of a ball rolling down a hill. No bottleneck at "can you visualize it."<p>So the narrower claim: it's not that intuition here is impossible full-stop, it's that human intuition is unreliable. Your post-hoc rationalization point is evidence for that, not against it. The story exists because a person needs something to hold in their head. Drop that requirement and the failure mode goes with it.
> A lot of people here are responding to the message but not to the meaning.<p>Well it is framed as quite specific <i>advice</i>.<p>(I'm done with mining PG tweets for meaning)
Couldn't have said it better
> understanding the bare-metal firmware for a computer<p>IMO this is still relevant, everything surrounding the LLMs needs such a vast infrastructure that I don't know if I would find it more useful to learn the maths behind ML than CS
As a 17 year old, I agree with this. Ofc I'm against all the hate directed at PG, I believe that all knowledge has value regardless of its economic utility, but I understand where the hate is coming from. Personally, I find LLMs boring for now, and I'm more focused on CS and electrical engineering.
Depends, the assumption things are predictable always negatively affects both Market Bears and Bulls alike.<p>Indeed, if credulous folks look to the world expecting people to bestow success upon them... than the disillusionment with reality will hit their savings harder.<p>The Shrek movie market correction correlations are undeniably funny, and a new film is due July 2027. OpenAI may be going public in the next few months while still losing $2.25 for every $1 of customer revenue, and with 6 other firms sharing over $4Tn in debt disclosed to investors in a footnote.<p>There is only one direction things can go at the Peak of inflated expectations. Popcorn ready. =3<p><a href="https://en.wikipedia.org/wiki/Gartner_hype_cycle" rel="nofollow">https://en.wikipedia.org/wiki/Gartner_hype_cycle</a>
> "Build an OS" wasn't a common university project because we were all expected to go […] helps you deeply understand how to intuit building for a whole class of problems.<p>Sure, but it was for a specific degree with a syllabus that taught you the foundational knowledge. It was not expected from the law students to learn how to build one.
> It would be a good idea for young people to deeply know how these programs work.<p>It would be a good idea for _everyone in the industry_ to deeply know how LLM training, inference and "agents" work, not least because it removes the ability of shysters to bamboozle with bullshit.<p>But, as much as a good idea it is for the young to understand this, it's the elderly who will be really taken advantage of if they do not keep up - just look at Facebook for good examples of why.
[dead]
There's this dilemma where in theory there's a ton of demand for engineers that can do real LLM machine-learning, but in practice there are very few available positions and entrepreneurship opportunities.<p>The reality is that an incredibly small minority of companies in the world do any real training or optimisation. It's unnecessary and inefficient for most purposes unless you are fully dedicated to being an LLM company, and still then it's a struggle. Those few that do train, they spend most of their budget on compute and have relatively small teams.<p>Getting experience in this field requires having access to very expensive hardware to begin with. And the skills will be quite hard to convert into any real value for someone, leading to a decent income, unless you have a ton of funding from patient investors, or you have decent contacts in Bay Area networks to get hired at the right place.<p>With all due respect, paulg is in somewhat of a bubble, this is not congruent with the global situation.
In that regard, it's not too different from mobile telephony. Mobile phones drove the electronics industry 20 years ago, but there is limited demand for people who <i>really</i> know how to build a phone (ie. build the hardware and write all the signal processing from scratch), as there aren't that many companies that do phones at the lowest level. A few of the engineers got rich (eg. Viterbi), but most 'just' made a good living. Most people who got rich off phones didn't do it by knowing how phones work.<p>Incidentally, the skills for the lowest levels of LLMs aren't that far removed from those needed for mobile telephony, in that both are based on maths, computation and information theory.
I was there at the start of the smartphone boom. I built a demo Android device that was capable of telephony/data, 3D rendering, etc. all before Google open-sourced the OS, for a SoC vendor that wasn't in Google's inner circle. Yet the industry was not interested in my junior profile during the subprime crisis.
This is super interesting because I moved from mobile telephony into ML and data science, and information theory and working with data in statistically correct way was what helped me! This was 10 years ago though.
Same with telecom companies.<p>We don't really have the demand for as many telecom companies as actually exist in the world. There's a reason we just have one Whatsapp and one Instagram, not three or four almost-but-not-quite clones in every single country that mostly differ in branding. The reason for the current situation has mostly to do with regulation and traditional, enterprise, "obviously every country needs a separate local branch, because that's what mcDonalds does" thinking. Technology has very little to do with it.<p>This is why the telecom world now consist of equipment manufacturers, who do most of the hard tech stuff, and actual telecom companies, who operate the equipment, rig towers in their local country, and maybe write some glue code to integrate a core from vendor A, a billing system from vendor B and a CRM / corporate invoicing system from government-approved local vendor C.<p>Banking also works similarly, though modern Neobanks / Fintechs and bank consolidation are slowly dissolving the concept of national bank branches.
Good points, but I think we can expect the AI space to be more tumultuous.<p>What most people want from mobile technology is for it to work, not too expensively, and for it to get out of their way.<p>What most people want out of AI is for no leader to emerge and wield supremacy against the rest of us. People are afraid of it in ways they weren't afraid of mobile, so they're more willing to work together against whoever is in the lead.<p>Its more like an arms race and less like a utility. The disadvantage I face when my competition has better mobile coverage and bandwidth is minor. The disadvantage I face when my competition has better intelligence on tap is much more significant.
Also, in a few years, LLMs will be building the next generation of LLMs anyway, probably autonomously to a high degree.
Yes. It's like looking at the (Apollo) moon rocket launch and then suggesting teenagers should learn to build rockets in their garages for the coming space age.<p>It is viable as a toy project, but there are vanishingly few career opportunities.
It's like looking at the early internet and then suggesting teenagers should write browsers as their projects instead of webpages.
If someone writes a browser in their teens, they will probably learn more about the web than if they were just writing web pages
They'll learn so much more that won't transfer to as many job opportunities. For ex, say more about C++ and less about cutting edge CSS (because modern browser tech is an ocean). I suppose they might luck into other adjacent or unrelated roles with the same skills.
I have a hard time imagining anything where writing a browser wouldn't be an excellent transferable foundation. Even your example: writing css will never teach you as much as writing a css engine.
CSS wasn’t around in the early web.<p>Building a browser in the early web was actually a very achievable goal for exactly the reasons it isn’t now. There was not JS. No CSS. No SVG. In fact very few widely supported image formats (and graphical browsers weren’t around in the earliest days of the web anyway). TLS didn’t exist. HTML only had a subset of methods. And even POST was usually just managed by CGI/BIN calling an external process, often written in C++ or Perl.<p>It was a simpler time.
They will also know the fundamental underpinnings of the web, you know - the thing that those html devs are actually using.<p>That’s like saying someone who fabricates cars doesn’t have the skills to drive them. Perhaps not, but they’re very well placed to pick it up quickly. They’ve also shown they can do something far more challenging, which is actually better than hiring for narrow immediate skills.<p>In other words: I’d hire that candidate in a heartbeat.
Are you asserting that C++ is not a marketable skill?<p>If anyone reading this has this un-marketable skill, Carnegie robotics in Pittsburgh is hiring<p>carnegie-robotics.breezy.hr/p/2d85f5321cc7-software-engineer
I think he's saying that building a browser is not a <i>transferrable</i> skill, like making a generic web page is. Employers don't like specialists. I used to build display drivers for graphics cards. Wonderful learning opportunity but other employers not in the graphics card manufacturing business didn't give a shit--those three years were essentially treated as an employment gap. "Well that's nice, but we really wish you had general experience writing CRUD apps..."
> I think he's saying that building a browser is not a transferrable skill, like making a generic web page is. Employers don't like specialists.<p>I recently switched roles, and among the seven places I interviewed, none of them seemed to see my then-current browser job as a problem, even though they were not related to browsers. (The closest one was a company implementing a HTTP reverse proxy, and I did not work on the browser's HTTP stack.)
CRUD didn’t exist in the early days of the web.<p>People keep talking about a browser in the modern context but the GP specifically said “early days of the internet” (which, in fairness, would mean pre-web. But I think it’s safe to assume they meant “web” not Internet).<p>In those days, it was actually a much simpler exercise to write a browser than it is today. I even wrote one! And writing a browser absolutely teaches you how HTTP and HTML worked. Plus a lot of backed development was forms data sent to CGI and thus written in languages we wouldn’t even dream of using for web development nowadays, including C++.<p>So in the early web, writing a browser absolutely was a transferable skill. It might not be now, but in the context defined by the GP, it was.
I would argue that in a post-LLM world, becoming as specialized as possible is the only way to survive. If you have the serious systems programming skills required to write graphics drivers, your talent would be wasted writing CRUD apps anyway. I hope you eventually found/will find something more appropriate for your skillset!
But building a web browser isn't a niche skill - it requires a whole bunch of them. You need to be a generalist to write a web browser.
This is like saying writing a compiler doesn't give you transferable skills. The surrounding competencies required to do this grant a pretty large amount of broad domain awareness
A few did.<p><pre><code> [Blake Ross] worked as an intern at Netscape at the age of 16 ... Ross became disenchanted with the browser he was working on and the direction given to it by America Online, which had recently purchased Netscape. Ross and Hyatt envisioned a smaller, easy-to-use browser that could have mass appeal, and Firefox was born from that idea ... in 2003 all of Mozilla's resources were devoted to the Firefox and Thunderbird projects. Released in November 2004, when Ross was 19, Firefox quickly grabbed market share ... with 100 million downloads in less than a year
</code></pre>
<a href="https://en.wikipedia.org/wiki/Blake_Ross" rel="nofollow">https://en.wikipedia.org/wiki/Blake_Ross</a>
In the meantime Mozilla was mocked on slashdot.org and elsewhere relentlessly after the first two years of the project when nobody believed there would be any value in the effort. Hats off to the team that took around five years to get to Firefox 1.0 (and released Mozilla browser in the interim). It took a lot of conviction to see it through.
Sure, you had to pick the browser I despise worse than Netscape 4.
That may have worked at the time, but no companies are interested in learning projects today. If you didn’t do the reqs list for the last 5-10 years with the same title, forget about it. Because there are dozens of folks who have, lined up. No one is indulging career changers (and most fresh grads) for the time being.
If you learnt to build a browser in 2000 you're probably doing well for yourself lol.<p>Not a lot of demand, but also probably not a lot of supply.
For this example, if you built a browser in 2000 it would put you in a great position to launch android in 2008. I think this is pg’s point. You spend all this time learning the interesting bits and when an inflection comes along that makes a new product or service possible you could be the one to likely launch it
That’s like saying if you can dunk in 8th grade you could be the next LeBron James. A LOT more things have to fall perfectly in place at the perfect time for that end result to materialize.
Each new bet will be expensive and depend less on technical merit than luck and having a backup so your family doesn't starve.
Agree.<p>FWIW I built browsers from 1999 for a long time. (But I was never a wunderkind, just somewhat tenaciously curious).<p>And I guess I am doing fine, but not amazingly rich or so.<p>Browsers were always a project closer to research/charity. I think Marc Andreessen said something similar - that he would never do that again. B2B is where you can make money.
I am talking from the perspective of the individual developer.<p>You don't need your impressive product to succeed to land a good career.<p>If Alice is doing LLM-from-scratch work today, and Bob is doing agent harness work today, Bob's project is far more likely than Alice's to become useful/popular/profitable.<p>But if neither project survives, in 5 years, Alice will be more employable/at a higher market rate than Bob.
This was his point.
Being prepare to launch a produce when the opportunity comes.
There is an entire graveyward of browsers, almost all that didn't die are now forgotten.
Sure, but I'm sure most of their lead devs are well paid now.<p>In general C++ work and similar, if not in Chrome development.
I think the point isn't that you'll necessarily build a great browser (or LLM), but the experience will benefit you in other ways.
If any were built by teenagers, I’d imagine those teenagers ended up with pretty good careers in technology?
Are you fingerpointing Marc Andreessen?
Well, there's a lot more to learn from the former than the latter ...
Really? The latter was immediately useful to lots of people which is motivating, and it had a nice smooth learning curve (html -> js -> php -> databases -> apps -> backend). Learning HTML is the first step to learning how to make full blown apps. Making a browser at 17 is like trying to climb Everest as your first hike. The expected outcome is burnout and demotivating failure. At best you'll learn some C++ or Rust.<p>17 is an interesting age. There are way too many comments here saying things like, 17 year olds should just do whatever seems interesting or bum around the world or focus on getting into university. But historically most kids were expected to be productive adults at 16 or 18. 17 is about the right time to be thinking seriously about what kind of work you'll do, how you'll make a living. University won't help and will just delay this decision.
Anecdotally, I think it's great advice.<p>I contributed to a browser engine around that age (KHTML, which later became WebKit and Blink), and while I don't work in browsers right now, much of that knowledge, mindset and of course the professional network have done much to shape my life. And a fairly successful career, for that matter.
Contributing to an open source project is fine, but the original analogy was "it's like telling teenagers to build browsers".<p>If teenagers could make small contributions to LLMs via open source then sure, go for it. Optimizing llama.cpp or similar would be a good learning project that might later get you good work via social networks. Contributing to open source is how I got started too.<p>Unfortunately, <i>training</i> LLMs isn't something that fits well to open source open collaboration. Inferencing codebases are better.
Excellent point, and I agree. Much of the benefit I saw was from working within a like-minded, smart team, not going it alone. And also specifically working on software with a real user audience to learn what providing value to them actually constitutes.<p>In that sense it's more a "seek out the open source community and real projects when young" rather than "do web browsers", with a bit of "look for ambitious types of projects few get to work on".
Historically people lived very different lives, required different skills, were poorer, had different opportunities etc.<p>A lot of people will make better decisions with a few yeas more maturity, and spending a few years developing themselves.<p>University will help a lot of people, and for some it will help.<p>There is a lot more to life than making a living.
Early web didn’t have JS. Nor PHP. In fact a lot of early web pages were written using static HTML with C++ invoked via CGI/bin for processing form data. So writing a browser would teach you the HTML plus C++ too.<p>The early Internet (which the GP mentioned) didn’t even have the web. But that’s nitpicking.
17 is a weird age but ymmv. I left home at 16 alone to study abroad. I had tons of free time due to dorm curfews and such. Unfortunately, I was too poor to have a computer and the computers we had access to were completely locked up. (Naturally we waltzed past the locks to play some games but it was also under surveillance)<p>Paradoxically I coded way more between ages 12-14, I regret my wasted late teens.
Me too:<p>ages 8-15: lots of great 8-bit computer fun.<p>ages 15-20: girls, booze, motorcycles.<p>20 onwards: get a PC, back to computers, realize how much I've been missing.<p>To be honest, judging by my own kids and their friends, late teens seem to generally be an era of hard to avoid stupidity.
I mean... that's actually amazing advice. Not because they would grow up to create browser startups. But because they would grow up to create web startups that succeed because of very fundamental of how the web is rendered.<p>Which is paulg's point really
This site exists because paulg created a webshop in a niche language (Lisp) and got a prototype bought out (and discarded) in the gold rush. No browser internals needed to make a fortune on the platform. Just rapid development of an application with natural monetization, and being in a place to do so way at the beginning.<p>Jeff Bezos did similar, but he did his own fulfillment and hired out the coding.
Early internet let me create the best personal web page in my city that I knew of with 2 weeks experience as a 12 year old. I imagine that same 12 year old could be more knowledgeable about LLMs than 99% of people in the same period.
You're vastly over-estimating the difficulty of making a web page and even more vastly under-estimating the difficulty of LLMs.
I doubt it. Creating a website can be done by copy pasting a few snippets together and checking if it visually looks like expected.<p>Good luck with that approach when trying to toy around with models and their training/inference.<p>There is also a lot of math basics missing that a 12 year old may be able to grasp, but I would bet they are at least 13 by the time the knowledge is deep enough to understand what operations are happening.
That seems like a perfectly good use of time for a kid in the 70’s! You’d learn a lot of engineering skills and demonstrate a tenaciousness that most people don’t have. I don’t think the goal is to predict what will be the important technology in 10 years. The goal is to challenge oneself with hard tasks and learn interesting stuff. A lot of the “AI” people today were compilers people yesterday
There are more career opportunities building rockets than designing rockets. Lots of welders, machinists, electrical engineers etc build rockets, and those skills transfer.<p>The same is not so different for AI. A few people design novel AI, but there are a lot of people training AI (especially if you include fine tuning) and implementing AI, even as a hobby.
Building an LLM covers a lot of CS fundamentals and forces you to do lots of research in order to implement one, especially on more modest hardware.
How far into the future can you see?
But he’s not saying that it’s a career opportunity.<p>It seems to me like he’s saying that doing this thing would be 1) fun and 2) a great way to become employable in the future. I don’t believe he’s saying that this project would be some kind of job training exercise.
When I was 17 and still witnessing Apollo Moon launches, I wanted to build an AI that would handily outperform LLMs as we know them today.<p>But that was way back in the early 1970's and all I had to work with was a mainframe.<p>Well the mainframe itself wasn't bad, the real show-stopper was that I didn't own the computer outright, no strings attached, no debt, etc.<p>>I'd probably try to make an LLM that I could use on some specific problem.<p>I thought so too back then, still do so I guess this is one of those things that could stand the test of time. I always wanted to start with something a lot simpler than a Moon mission myself. At 17 I already had a significant breakthrough in the chem labs and it was from alternatives to a single processing step plus everything that descended from that, rather than trying to tackle a much more complex detailed multi-step synthesis. I was only 17 but I was not trying to be a slouch, I don't think pg was either at that age but his advice is not for just anybody. I couldn't have done it if I hadn't made major progress since being 16, and it really emphasized at the time how much maturity can make a difference. My imagination ran wild as I extrapolated :)<p>In a reply from LeCun to pg:<p>>>I'll figure out a set of methods and architectures beyond LLMs that can quickly learn to perform physical tasks as efficiently as humans and animals.
That last item is also what I would if I were 30, 40, 50, or 66 years old<p>I see no reason to stop at 66 either ;)<p>But I figured that people owning more computer power than I could ever afford were going to be doing something like this as soon as they could, without having to wait for something like an LLM to arrive before getting peoples' attention.<p>It did seem like things were going to take longer than you expect, so it's pretty good to have a lifetime of concentrating on the specialized natural science domain expertise, focused now for 50 full years on how it would combine if AI ever got good enough.<p>Both the natural science and the AI need to be a major cut above, I still see dramatic room for improvement in my own work. If I'm going to have to rely on "other peoples' AI" then that natural science component is going to have to pull a lot of weight to keep up with the kind of computers that only rich-as-hell high-rollers have access to.
I agree with mostly all of this, but personally I wrote a toy LLM almost 5 years ago and while it never saw much use outside of boring my wife with a shitty command line demo with glee it did help me understand how they worked and how to apply them, played a lot with JAX and pytorch, ended up building a ghetto version of MCP and an LLM-Pool to proxy requests to my baby local models and so I didn't struggle to see the evolution of openrouter and MCP agentic workflows. The same way i'm really glad when I was younger I built a bad webserver by myself, a really painful SQLx type database, etc etc etc - none of these things led me to developing for Nginx or Oracle nor will knowing JAX get me a job at an AI research lab, but I do have a lot of depth in understanding how the technology works so that the flavors on top of them are easy to digest and make more use of immediately, and I think the same can be said for engineers coming into the field - if it's a spooky LLM box you aren't going to be squeezing the same amount of juice as the guy that knows how they work inside and out so having at least the understanding of a _babys first LLM_ is going to get you miles ahead of people who don't.<p>For anyone who wants to dork around there is <a href="https://github.com/rasbt/LLMs-from-scratch" rel="nofollow">https://github.com/rasbt/LLMs-from-scratch</a> which is something amazing that I think anyone who wants to engineer things around LLMs should at least blast through and read.
Game cheating and reverse engineering MMO backends taught me a lot: databases, networking, securing a backend (and frontend), limitations of simpler languages when comparing them to more native options for building backends.
Agreed, I was very late to the game and was forced to learn VBA for excel sheets and that is how I finally broke into programming.<p>When I was a pre-teen I stumbled upon CD-rom hacking guide to bypass disc requirements on games, I remember opening up the file and the screen being filled with HEX code. I was so overwhelmed I just closed it and never touched programming after that for 15 years. My life would have been totally different if I had embraced the unknown instead of retreating.
This was my introduction, too, but with Counter-Strike cheats.
It was Quake World for me
> The reality is that an incredibly small minority of companies in the world do any real training or optimisation. It's unnecessary and inefficient for most purposes unless you are fully dedicated to being an LLM company, and still then it's a struggle. Those few that do train, they spend most of their budget on compute and have relatively small teams.<p>And the job postings are often ridiculous. I recently was an AMD job advert in Germany for an ML Kernel Engineer, not Senior mind you. The requirements went something like<p>> Masters Degree required with strong preference for a PhD with peer reviewed articles in {journals_list}
> 10+ years of experience in C/C++
> GPU programming experience required
> 10 more ridiculous lines<p>No idea how a teenager self teaching himself LLMs is supposed to even get a shot...
It reminds me of that "stone soup" story.<p>1. I can make turn a stone and water into a delicious soup<p>"17 year olds, learn to build an LLM from scratch"<p>2. This soup would be more delicious if we add a few carrots. Does anyone have carrots<p>"Increase your chance of success by getting a Masters degree"<p>3. How about potatoes?<p>"And get a PHD"<p>4. What about some salt?<p>"And publish some peer reviewed articles in {journals_list}<p>5. We should also add beef<p>"Now work in the industry for 10 years"<p>6. See, this soup is delicious, and I made it all with a stone<p>"See, you're rich, and it's all because you learned LLMs as a 17 year old"
I have always been pro fundamentals. It caused me trouble early in my career with bosses that didn’t understand why I would spend time trying to understand how something worked at a low level if I was a high level user. But then knowing the fundamentals gave me an edge as a designer and developer by understanding capabilities and limitations of the tools I was using. For example understanding how indexes work internally in a relational database. So I see the value in this type of work, not to land a job as a LLM researcher, but as an informed user of the tool.
I’ve been wondering about this.<p>There are high school students competing in contests that cover parts of the (Math) theory behind AI. A lot of high school research programs are integrating AI with other things and complex mathematical models…<p>To me, this is bizarre as Calculus is barely taught in high schools (and likely poorly).<p>Don’t get me wrong, these kids certainly aren’t the usual lot.<p>Yet, I really wonder if they know the fundamentals. Do they even understand derivatives or just memorized the rule for polynomials? Can they even explain what a transistor is?<p>Normal curriculum takes 5 years to go from Algebra I to Calculus. Real Analysis, Linear Systems, etc. are fundamentals taught only in college…<p>Feels like too many are trying sprint before even learning to walk.
I disagree with the premise.<p>Learning should not be done only as a direct path to getting paid.<p>Learn to create pattern matching and intuition to solve future problems.<p>When you are 17 it is a good time to understand how the world works so you can build on top of it in the future. If we assume most tech is going to have an LLM as part the stack, a solid basis in how LLMs work is likely to help you in future endeavors the same way a solid basis in how the web works helps you today.<p>Maybe a 17 year old should learn both. As a small anecdote when I was 17 I learned a lot about load balancers, failover, and building self-healing systems running small hosting company that had to be fault tolerant when I was attending high school. This wasn't at state of the art levels (e.g. I wasn't configuring gigabit routers or global CDNs -- but it was useful pattern matching for future problems)<p>I currently don't touch any of that tech, but I have working knowledge that still serves me today.<p>Think long term.
Yes. There's a lot of demand for elite talent, and no demand for slightly sub-elite talent.
AI is the subtrate the future runs on.<p>And so I think the idea is more to understand tomorrow ... from first principles.<p>In the late 80's, as a teenager, I learned x86 assembly and C because that was the only way to squeeze out enough juice from my shitty CGA (and later VGA) card to programm the games/graphics that interested me.<p>I haven't written assembly in years.<p>But whatever I did in my career: it helped me and gave me an edge over my peers to have a foundation that is very close to the metal.
> AI is the subtrate the future runs on.<p>We don't know this. So many people are simply claiming this confidently, and a lot of them are betting their careers on it, but nobody has a crystal ball. Whenever someone tells you confidently, and without any doubt or qualifications, that something "is the future," be skeptical.<p>I remember when the Segway was definitely going to change urban planning worldwide.
> <i>AI is the subtrate the future runs on.</i><p>Current AI can automate significant amounts of grunt work in programming and math. It's good at running web searches and writing summaries. There are a few other niches where it is currently successful. But other than that, many corporate AI projects are spectacular failures.<p>So just given what we have in hand, assuming no further breakthroughs, then we're maybe looking at AI being somewhat bigger than the Internet. Which would make it a revolutionary technology, sure.<p>But to get from "a revolutionary technology" to "the substrate the future runs on", then you need to assume more breakthroughs: long-context operation over weeks or months, displacing human workers 100% instead of 75%, and the ability to directly economically compete with actual humans. And people are investing literal trillions of dollars to make that future come true, without really thinking through what truly competitive-with-human AI would actually mean. We might be looking at massive job loss, centralization of power, fully automated "companies" with no humans dominating markets, and other dystopian scenarios.<p>And in those worlds, it's unclear that being good at CUDA and matrix math will be all that helpful, careerwise. The AIs are <i>already</i> pretty good at that stuff. Data scientists get paid OK when they actually get hired, but it's not everything college students were promised in the 2010s, either.<p>We can't yet build a fully-general competitor for the human mind. But we're getting closer. And if we ever do build one, the consequences will be really weird in any number of ways. So I worry about visions of the future that assume AI keeps improving significantly, but that also assume it still somehow remains a "normal" technology that doesn't, for example, render most humans fundamentally uncompetitive.
> But to get from "a revolutionary technology" to "the substrate the future runs on", then you need to assume more breakthroughs: long-context operation over weeks or months, displacing human workers 100% instead of 75%, and the ability to directly economically compete with actual humans ... without really thinking through what truly competitive-with-human AI would actually mean. We might be looking at massive job loss, centralization of power, fully automated "companies" with no humans dominating markets, and other dystopian scenarios.<p>At 75% replacement of a worker we would already have huge job losses as each individual would be doing what several before did.<p>The only alleviation would be the creation of new equivalently paid jobs, which is no better than a hypothesis right now.
>> many corporate AI projects are spectacular failures.<p>Citation needed
I keep saying AI is going to prove more impactful than cloud computing but less impactful than the sewing machine and a bunch of people get mad at me.
I don't really agree, I think the future should run on humans.<p>AI is a great solution in niche areas but generally doesn't make much money. All the large companies are in the negative.<p>The steam engine was less of a bubble and was much more revolutionary and had a greater impact.
How did that edge manifest?
That's like saying the only way to do real engineering is with Google-scale Borg deployments. You can do quite a lot on very little hardware, r/StableDiffusion is a prime example.
You can do plenty of "real engineering" under normal conditions. But specifically when it comes to LLM engineering, no there's really not much you can do, they are called "large" for a reason. You can play around at small scale, but those lessons you learn will not be very relevant to the real problems in the market.<p>Sure you can gradually climb the ladder by demonstrating your skills bit by bit and getting access to more resources. It has very good prospects if you do manage to push through. But it's a hard and risky path, and you will not be able to get any interesting results for the longest time.<p>For a young middle-class student, it just doesn't make much sense. You can do much more impressive and impactful things with your time without getting into that black hole.<p>I know how to build an LLM, I know plenty of fellow young engineers that do too. It's really not that complex. But they can't do much with it without capital or access.<p>Good engineering has never been a bottleneck in this field, it's been all about having access to capital and taking smart but dangerous risks burning it on compute, without much idea of how long you need to keep burning for. There's still no end in sight, some are still managing to convince investors and keep burning, and we are seeing progress, but the business case is still unclear. If you want to get in that game, go ahead, but it's not something I would advice the average young engineer.
> I know how to build an LLM, I know plenty of fellow young engineers that do too. It's really not that complex.<p>I’m 40, and I don’t.I took that abstraction for granted and “left it to the big labs”. However I want to build my own LLM for learning purposes.<p>On needing big expensive hardware.. necessity is the mother of great innovation. Perhaps 18year olds trying to build their own LLMs in constrained resources environments will result in ground breaking ideas of achieving better intelligence than the one we currently have….<p>The world needs pragmatic folks who work at a higher abstraction and make LLMs useful, AND also folks who think why not “this other way”? And build newer ways to do fundamental things.<p>Given the usefulness of current LLMs, I would certainly encourage anybody to try and build their own LLMs, and see what they come up with…<p>Heck if they build a rack full of old laptops and run something with it that could be done “better” with modern servers, I’d still appreciate the learning running things on those little machines bring.
Well, that's not how it works. You don't just put some old laptops into a rack.<p>Maybe with a decent consumer GPU like a 4090, you could do experiments like distilling and fine tuning a small image model for edge deployment for specific tasks.<p>Even there, many use cases might require renting compute for $10/hour and investing a few hundred.<p>A LLM from scratch? Forget it. You can do theoretical experiments, but not build anything remotely useful with that kind of budget.<p>If you're talented enough to come up with revolutionary methods, maybe an university or AI lab would be the place to be.
To me the bottleneck is not even the compute, which is an issue for sure, but the data. All these large companies got their hands into petabytes of data, <i>a lot</i> of which of illegally acquired, but now they are large enough to pay the fines.
> On needing big expensive hardware.. necessity is the mother of great innovation. Perhaps 18year olds trying to build their own LLMs in constrained resources environments will result in ground breaking ideas of achieving better intelligence than the one we currently have….<p>10000%.
One of the things to think about when it comes to many kinds of "expensive" technology, is that from so many well-funded ventures, from the capitalists on down almost every decision-maker involved has not spent the majority of their life making <i>every dollar count</i> in some way or another.<p>Even more so when things are not just expensive by nature, but truly <i>overpriced</i> beyond that point.<p>>10000%<p>Once in a while you do get somebody who only spends a dollar and gets more out of it than a seasoned high-roller spending $10000. Most of the time the waste is borne by those who can afford to throw away $10000 more easily than an economizer can afford to lose one dollar, so nobody is crying about it.<p>With how ridiculously large the language models have gotten though, a 10000x improvement in actual intelligence does seem like it could be lurking unrecognized at a different point on the compass.
Agreed. It's hard to learn unless you have access to quite high end hardware, and even paying by the hour is expensive. There's a low ceiling on what you can learn without doing training runs.<p>You can however learn everything you need to know to get on the career ladder as a software engineer on a regular home PC.
While the topic here is narrow, the concept is broader.<p>Do you take the first step or rule it out because you don’t yet see the complete picture.<p>As a teenager I never hesitated to try things out. As a young adult I wanted the whole picture. Now I’m back to playing / trying things out. I kinda wish I’d not given it up. PG being a bit older and reminiscing - I bet he’s in that bucket too, whereas someone trying to establish themselves professionally probably (aka me early 20s) wants to see the path.
> . But specifically when it comes to LLM engineering, no there's really not much you can do, they are called "large"<p>The “large” qualifier dates back to pre-transformer language models, where even training a multi-million model was hard due to how poorly it scaled. GPT-2 was a <i>large</i> language model, despite being only 124 millions parameters.<p>Due to how much high quality data is readily available, anyone can now train a sub-billion (L?)LM on commodity hardware.<p>And I'm personally convinced that pretty much any enterprise use-case of an LLM (except coding) is better served by a fine-tuned small (<2B) model that is trained specifically on the task, rather than a generalist frontier model, so learning the engineering around fine-tuning is a key skill that companies will realize they need sooner than later.
Why? Oersted is correct, for any size class you can find an LLM that is free and well trained at this point. They are highly adaptable even without fine tuning, in-context learning is still superior to fine tuning in most cases also. And real world fine tuning is mostly about data gathering and cleaning. The actual adapter training is automated and put behind simple APIs.<p>I saw my first language model in action in 2014, I was writing blog posts about them back in 2016. In recent years, like many of us, I spent some time learning ML frameworks to see if it'd be a fun career pivot.<p>But:<p>1. It doesn't seem especially creative. All the frontier labs have nearly identical model personalities, capabilities and even app designs. We're not seeing them differentiate from each other, implying that the design space might not be that large. In which case the opinions and unique approaches of specific engineers aren't that important, they are interchangeable at the right level of skill, and what to do next is usually obvious to everyone.<p>2. It doesn't seem like a big job market. A lot of ML jobs were wiped out in recent years by the rise of LLMs. Lots of NLP specialists etc were suddenly replaceable with a cheap API. The jobs that remain have compacted into a small number of companies. It's a small community which greatly increases career risk, especially as so many are unprofitable and/or have strong ideological requirements.<p>3. It's unclear how much demand for better models there actually is. Do we actually need smarter models? In robotics clearly yes and robotics <i>is</i> interesting and high potential, but for pure LLMs/image models, most users are already incapable of setting tasks that stress the best models and are happy with the cheaper smaller ones.<p>Using the models on the other hand is a very large design space, and has a lot of scope for creativity. I see use cases for AI everywhere, but most companies seem to stop at putting a chatbot on their website or asking Copilot to rewrite an email before they send it. A lot of companies have hollowed out their IT departments over the past twenty years. It feels like a new golden age of consulting work could be upon us.<p>I really don't think I'd tell a 17 year old to learn how to <i>train</i> LLMs. Learn how they work and how to use them, sure, absolutely.
> We're not seeing them differentiate from each other, implying that the design space might not be that large.<p>There is the more likely reason they are not differentiating. They use almost exactly the same class of model. Everything is linear, parallelizable. It's incredible path dependence that's now invisible enough we think it's a natural law. Nature is not linear.
> They are highly adaptable even without fine tuning, in-context learning is still superior to fine tuning in most cases also.<p>Good luck relying on in-context learning for a 600M LLM.<p>> The actual adapter training is automated and put behind simple APIs.<p>That's like saying it's worthless to learn infra because you can use serverless instead…<p>> All the frontier labs have nearly identical model personalities, capabilities and even app designs. We're not seeing them differentiate from each other, implying that the design space might not be that large<p>The design space for a <i>generalist</i> model isn't large, by definition. But the design space for specialized smaller models is much larger. If you can train a 200M model that, for your use-case, is competitive with a frontier one, then you'll make your company save a lot of money in tokens.<p>> 2. It doesn't seem like a big job market. A lot of ML jobs were wiped out in recent years by the rise of LLMs. Lots of NLP specialists etc were suddenly replaceable with a cheap API. The jobs that remain have compacted into a small number of companies.<p>We are in a strange place where a few companies are collectively burning a hundreds of billions a year to sell things a few pennies for the dollar. Of course it's going to be cheap and concentrated. How is it supposed to end though?<p>> 3. It's unclear how much demand for better models there actually is. Do we actually need smarter models?<p>That's the thing actually: I don't think we need better models this much, and if we don't need better models we need the cheapest possible model for a given use-case.
But can you _sell_ that? If you can't you can't get a job doing it.
Interesting/capable diffusion models are much smaller than similarly interesting language models. But yes you could always scale things down to learn the fundamentals.
A lot of startup companies are not training frontier models but help solve and optimize pain points of LLMs: cyber security, token usage, harnesses etc. These jobs don't require a PHD in machine learning but it does help if you understand LLMs at a deeper level.
A single 3090 will train qwen 0.8B just fine. While it’s not a very capable model any training technique you would want to master can be used to make real progress. And all the skills you need to learn how to do this can be learned watching Andrej Karpathy’s zero to hero series (shame he quit educational content and went to anthropic)
But you can train a small LLM with a gaming graphics card -- I managed one on a GTX 1660. I don't think pg is suggesting that you try to chase the frontier. It's more like building your own OS in the 80s, or web server in the 90s -- sure, you'll never match the commercial offerings or the big OS projects, but building something from scratch within the limits of the hardware you can afford is amazing educationally.
when you think about all of the advancements since Attention / GPT a lot of it has been somewhat more obvious than in other fields, as is typical with the massive flood of innovation that follows a big breakthrough.<p>Paul likely assumes there will be a sequence of additional papers with the same impact as Attention is all you need, which will spawn a lot of opportunity for a larger group of experts who are conversant enough to advance the field even if they do not themselves create such a major innovation. Not only is this deeply exciting, it is also highly meritocratic as there is still scarcity of the kind of intellect and creativity necessary to swim there.<p>Machine intelligence might soon surpass it, though, and Deepseek is 100% Chinese mainland educated. Paul's description of building an LLM from scratch is meant as a vague starting point for being an innovator of the highest value aspect of modern AI innovation, not as a specific prescription.
Y'all are missing the point: It's probably less than 100 hours to learn the foundations of one of humanity's most-current breakthroughs. It's a disservice to any young hacker to <i>not</i> learn it. Here's your curriculum. Watch these:<p>-- 3Blue 1Brown's Neural Network Series: <a href="https://www.youtube.com/playlist?list=PLZHQObOWTQDNU6R1_6700" rel="nofollow">https://www.youtube.com/playlist?list=PLZHQObOWTQDNU6R1_6700</a>...<p>-- Karpathy on LLMs: <a href="https://www.youtube.com/watch?v=7xTGNNLPyMI" rel="nofollow">https://www.youtube.com/watch?v=7xTGNNLPyMI</a><p>-- Stanford CS336: <a href="https://www.youtube.com/watch?v=JuoVZkPBiKk" rel="nofollow">https://www.youtube.com/watch?v=JuoVZkPBiKk</a><p>Then do this hands-on:<p>-- Karpathy's zero-to-hero: <a href="https://karpathy.ai/zero-to-hero.html" rel="nofollow">https://karpathy.ai/zero-to-hero.html</a>
This is a wildly incorrect and myopic view on the world.<p>Finetuning model is cheap and incredibly useful for deployment. You don't need to pre-train a frontier llm from scratch to make useful models.<p>There is tons of domains where you and fine-tune llms and deploy them for value in companies and for your own entrepreneurship ambitions. I have made this a big part of my career for the last few years and now I'm working on finetuning models for starting my own companies.
I find the fine tune approach more interesting than straight to RAG and MCP.<p>End of the day they're all customized data stores and protocols to interact with them. May as well stick to a uniform toolkit with fine-tunes.<p>Not that other tools aren't useful. But reaching straight for a bunch of infrastructure reliant services is like jumping in with k8s when you're still at a stage where basic mocks in code are sufficient.<p>I won't roll my own encryption or UI lib but want to stay focused on the incompleteness of the project I have to ship not all the buttons and knobs of some dependency or framework. Same old manage context switch problem.
Both of you are right. There is demand for tailored (fine-tuned) models; almost every enterprise would theoretically benefit from them.<p>But there are also a lot of prerequisites, namely does the enterprise have its sh*t together on a technical level. Does it have the processes and data pipelines available to train and benefit from these models? Probably not!<p>Applied ML is at the crown of a tech pyramid whereas most enterprises are still struggling at ground level. Being able to build from be ground is likely a safer skillset than only knowing how to work at the (non-existent) apex.
I think his point is to do this to understand deeply what they can do, what they can’t do, and what they can almost do. And then find the highest value ‘almost’ use case and push there. Which doesn’t necessarily mean improve the llm, could be applying it in just the right way for the use case. Of course, the bitter lesson makes this hard and risky. But no more risky than investing your time in learning anything else these days.
I bet this will get less true over time though as the rate of change slows down, allowing specialized models/training for specific use cases that aren't TAM heavy enough for the big labs to go after them. It's just now any general model is the best thing to use for everything and you're wasting money to build something on what will certainly be obsolete by the time you can get it to market
Yes and no. I believe the point he is making is simply that there is no substitute for fundamentals and first-principles thinking.<p>We had scores of students study how microprocessors work and compilers work over decades, yet we have 3 or 4 major processor companies and a handful of programming languages. Yet, what they learned was probably crucial in their development as engineers.<p>We are also so early right now that even 2-3 years from now who knows how many LLMs and model firms survive (esp. given the "snake eating its tail" venture/investor funding situation)
This is kind of different though isn't it? Doing an assembly or compiler class has pretty clear benefits in this regard.<p>But LLMs are tools. Does a great engineer need to know how vscode works? Might be helpful to understand how extensions work, LSPs, and project configurations.<p>Usually when working with any tools, you need to understand how to get the most out of your tool for your needs and that's about it. Core fundamentals about how software and hardware works in general seems like it would be MUCH more useful than LLM core knowledge.
I don't think pg is giving advice on what will lead most directly to a job, but rather what is the best learning for a 17yo.<p>A 17yo who trains their own LLM will have a much richer understanding of what AI is, how it works, what its potential capabilities and pitfalls are, versus someone who spends the same time doing something else.
We finetune LLMs. Small ones like Gemma 4 for semantic tasks.<p>There are plenty of areas were we need people to do this for insurances, banks etc.<p>AI/ML exists on many levels.
"Necessity is the mother of invention" - limited hardware has always forced people to find cleverer ways of doing more with less. Current models are clearly nowhere near the efficiency limit (the brain does vastly more with far less power).
This is roughly how GPUs for neural networks got started: after Andrew Ng left Google Brain, he no longer had access to a 10,000-CPU cluster used to train the original DistBelief system. But his Stanford students could buy a GPU...
> Current models are clearly nowhere near the efficiency limit (the brain does vastly more with far less power).<p>I think this is disingenuous. One could say that drones are nowhere the efficiency limit either: a bee can fly for hours on the energy contained in just a few milligrams of honey, while our best battery-powered drones can't stay airborne for more than 30 minutes. But comparing energy efficiency of electric/mechanical devices to their biological counterparts is not an apples-to-apples comparison. There's a world of difference between the energy storage and delivery mechanisms.<p>And as many have pointed out already in the siblings, it's not just about the compute but the access to petabytes of training data.
It is not a skill that you will use in your day to day life, but I think it is part of the fundamentals now. Sure, LLMs are in a bubble, just like the web during the dotcom bubble, but web didn't disappear, and I don't expect LLMs to, even after the bubble bursts.<p>I didn't write a LLM from scratch but it is on my "wishlist" so to speak. From what it seems, a GPT-1 class LLM can be done from scratch in a few days and tens of dollars of cloud compute or a high-end gaming GPU.<p>It is an exercise not unlike building a compiler, a school classic. You are unlikely to ever work on a compiler, but at least, now, you know your tools a little better. It is not about becoming an expert, that takes years, it is about knowing what you are doing.<p>If you intend to make software engineering your career, you will want more than surface knowledge. And that part is entirely on you, or on your school if you are a student. Companies will not pay for you to learn the fundamentals, they want short term returns, because you may leave at any time. But you as a software engineer may have 40+ years left, so it is worth thinking long term. Claude code may become obsolete a few years, but linear algebra is not going anywhere.
Of course, because it's not the LLM that's special but the training data. Nowadays, your favourite AI service to generate code for an LLM whenever you ask for it.
I don’t think that’s correct, the data is not that special either, and getting a similar dataset is significantly easier than getting the compute capacity to use it, even if they are both relatively hard.<p>Probably this also is too cynical and simplistic, but: really what’s special is the ability to get this kind of capital, with the freedom to burn it on mad moonshots, with long enough leeway to actually get to see a few of the moonshots come true.<p>No wonder that the head of YC made this happen, this is exactly what they are world-leading at.
You can train and run small models on an old gpu. That’s what I’m doing now at, well, much older than 17. Does it produce a useful model? No. Not even remotely.<p>However, I do learn stuff about models that takes it from “magic” to “useful tool I understand the limitations of.”<p>Do I do it for that reason? No not really, I’ve never had luck learning something because it would be good for my career. I do it because at my core I’m a bored teenager who wants to make the computer do cool shit.
It is on my list to build a toy LLM from scratch.<p>Not that expect to make it big as a LLM researcher but building something from scratch gives a much deeper understanding than what you can get from simply using something.<p>Much in the same way as implementing and designing your own programming language makes you a much better programmer.
This is a silly take. You can learn to build an LLM, there are great resources to do so (there are books about building them from scratch), you can use older model GPUs or rent them by the hour. The value of understanding them is really high for anyone building any application that uses an LLM at any point.<p>It similar to understanding how a very basic CPU works. Just because I'm not going to work at intel or nvidia or whatever optimizing the hell out of a chip, it doesn't mean I just throw my hands up and think "magic" - the basic architecture isn't that difficult, and the value of knowing it is astronomical for anyone writing software.
What about other machine learning related skills? Does this wave of LLM mean less need for that kind of work?<p>I would think not, but when I started to look into OCR options recently - assuming that obviously a dedicated tool would do a better job than an LLM - I was wrong (apparently).
> With all due respect, paulg is in somewhat of a bubble<p>I feel like the "ALWAYS HAS BEEN" meme is apropos here.
+1, I have a friend, math PhD that's been working on ML research 5+ years in London yet he has not been able to find any position.<p>The only jobs that he found he was highly over qualified or paid very little.<p>In any case, it doesn't look like there's this crazy rush to hire all ML talent, even the one that understand the math and technology deeply.
> +1, I have a friend, math PhD that's been working on ML research 5+ years in London yet he has not been able to find any position.<p>Maybe people simply don't want math PhDs but something else? Since 1-2 years ago I started doing consulting/freelancing in the ML space, but more on the infrastructure, deployments and similar stuff, as a general purpose developer, and I have a waiting list of clients interested in more work, some of them even trying to recruit me to work for them full-time as well. I'm based in continental Europe, fwiw.
How've you gone about getting into this btw? I have extensive experience in infra and pipeline rollout but have struggled to find freelance clients for this kind of thing. Would be great to tie it into ML as a learning opportunity there
Spent a year of freetime catching up on everything and learning as much as possible, started sharing what I've found works or not, write a bunch of comments on HN and elsewhere, and have a email in your profile, eventually people will find you if you put out good stuff :)<p>Also bunch of past workplaces who've adopted AI in various ways who reach out once they find out what my current focus lies, but that's harder for others to replicate unless you've already had a career as a developer.
whats ur contact, would love to chat
This only proves the original point which is that there is not much demand for actual machine learning expertise because that is only carried out in a small number of places and what demands there is is for the more basic software carpentry like infrastructure and operations rather than the actual technology and Engineering side of things
What parent says about "there are very few available positions" for "engineers that can do real LLM machine-learning" is fair, yeah, I'd agree with this.<p>I don't think the "incredibly small minority of companies in the world do any real training or optimisation" part is necessarily as true, as some parts of the work I do get is about helping them optimize training and infrastructure around training. Mind you, none of this is for building LLMs from scratch, it's 99% fine-tuning existing checkpoints.<p>I'd also agree with "paulg is in somewhat of a bubble" regardless of this, which is worth remembering whenever you read his content. Same goes for any person living in SF, and dare I say the US. But also, YMMV, I live and work in Europe, probably why I have this perspective.
Congrats!<p>Is it possible to see some of your old works? Personal research?
As always and everywhere, it's who you know (and who knows you) that matters.
I mean it makes sense right? Anthropic for instance has like, a couple hundred staff in London with plans to expand to somewhere just shy of a thousand. There are far far more ML/Maths/CS/Stats PhD's than there are openings. Especially in London there is no shortage of suitable candidates given Cambridge/Oxford/Imperial/UCL are surrounding it. 2% of the UK population has a PhD alone...
I think a lot of this is based on preconceptions. A lot of apps were made with Electron, because it was common wisdom that native is 'too hard'.<p>Now with LLMs, people write native apps in Rust, and I'd like to think some of them found that there isn't such a huge jump in difficulty they assumed there would be.
It never was native "too hard" it was always "too expensive".<p>That's the same case finding companies that will actually pay for hand made LLM instead of using something from big providers will be hard because most companies won't be able to afford it.<p>Yeah if you find a company that will do that stuff directly, good for you, but you will have to be very lucky and you will have to compete with other people who followed PG advice.<p>So I would rather learn all there is about properly using LLMs and integrating them with existing systems, that will most likely by useful for 90% of companies out there.<p>Building business niche harnesses is in my opinion much better direction. Knowing what will work best in specific cases is it FTS or vector search, optimisation of usage, getting best results while using cheaper models, knowing how to use tools to run models on the servers, and all the tooling around that like various MCP or just tooling that will be provided to models.<p>That is what I am currently busy with and I already have customers for that knowledge.
Oh yes I agree, LLMs are not that complex in principle, most engineers could build a toy version completely from scratch without too much difficulty.<p>But that’s the tip of the iceberg. If you have any ambitions of doing this professionally, it quickly becomes clear that all it’s all about knowing how to deal with problems that are only present at massive scale, when an LLM is actually L and becomes AI.<p>The mundane details about how to build a tiny autocomplete model and the maths behind it you can learn in a couple weeks easily. It’s not black magic, there are much harder areas of computers science.
FWIW, at least 20 Y Combinator startups have published ML research recently at ICLR, NeurIPS, ICML, and so on.<p>I think a lot of people assume that only the big AI labs can do cutting edge research, but there's a strong argument you can do it as part of little tech as well.
The other thing to add as well is that the research teams who do the actual research work are relatively small and very specifically qualified which naturally keeps the barrier to entry high.
lol unnecessary and inefficient…<p>I just finished fine tuning Gemma e2b for local code completion on my local machine.<p>This comment just reinforces what the post actually means. We need people that are LLM natives, computing solves itself with time and with scale adjustments
That’s great but the industry has moved on to code generation and editing, not code completion.
1. I very much still write code myself without an LLM when I need top quality<p>2. That's why I have an agentic agent as well installed, Qwen 27B, outrageously good, better than sonnet 2 years ago. And it's mine, I can give it confidential info to work with since I own the whole chain. See where I'm going with this?
Can you tell me what specs your machine has? There is a difference between a few hours and a few days
It is written in the first person, I suppose.
I think companies of all sizes will want their own models, or at least customised ones, for their own specific use cases or competition and security issues.<p>1. Both training and optimisation will get significantly cheaper and easier quickly.<p>2. Politics will probably get even more insane before a potential reprieve on the 20th of Jan 2029.<p>3. The big AI firms will become part of the surveillance capitalism network, if they're not already.<p>So I think for self-protection a lot of companies will be looking near to medium term AI independence.
The argument is sound, but the maths don't math for now, and it's unclear when/if they will.<p>For the time being, unless you truly have millions, the outcome from training will be very net negative, while focusing on building on top of existing AI will yield amazing things if you apply the same talent and effort.<p>When it does get cheaper, then it will be easier to acquire the skills and experience too, and the struggle you went through by trying to do it now will be somewhat wasted.<p>Besides, I am well versed in this field, and it is not rocket science. There are plenty of software engineering domains that are a lot more challenging, like high-end graphics, large-scale data engineering or kernel programming. People will learn to train LLMs when people want them to.
Right, just like companies don't use SAAS.<p>In reality, enterprises are happy to offload even risky tasks to others as long as they get some contractual guarantees about their data. Would they like more choice in who to buy from? Yes, but not enough to in-house such a specific discipline.
The cost of training a model from scratch is going to be cost prohibitive for the vast majority of companies (even if renting the hardware needed for the 1-2 month training time). It's an interesting learning exercise, and some of the things learned can be applied to other parts of the process. There's also the issue of needing a huge amount of data needed to get decent weights.<p>Fine-tuning a model or LoRA based on the companies data set is more feasible but you're likely going to need several runs as you test/try out different base models, parameters, etc. This is why there are a lot of fine-tuned models on huggingface based on base or instruction-trained models from the larger AI companies that have released open weight models (Microsoft, Google, IBM, Mistral, DeepSeek, Qwen, etc.).<p>Training is limited on memory first (storing training data and weights) and computation second. Realistically you need to own or rent 2-8 H100/B100 devices or Google's TPUs.<p>The majority of workflows for a company providing AI capabilities are likely best solved by tailoring a system prompt for the chosen model, evaluating the prompt and model with tools like promptfoo, and then running it on a compute cloud provider (including AWS Bedrock). If the company is big/financially well off enough they could look at buying the hardware needed to run it on their own servers.<p>For other uses like agentic software development you'd need to spin up a suitable model on a compute cloud provider (or local hardware if the model is small enough) and then tell your IDE/editor to use that model. You would need some way of benchmarking and evaluating the models to see if they are capable of doing the tasks you need. -- There have been some tests done by people on YouTube that suggests that Qwen 3.8 27B is a decent model, but your needs may vary.
Even for most organizations, testing AI systems is too cost prohibitive, so they YOLO in production, including public facing systems.
Most companies that build physical goods don't care for one second about their IT department other than how much money they can save per month, starting by outsourcing whole of it, thus they have little use for internal LLMs.
And it's across the industry, thinking banks, private banks, insurance, pharamcy etc don't outsource their IT, including development...
I believe US outsource even more than Europe on this matter.
> <i>With all due respect, paulg is in somewhat of a bubble, this is not congruent with the global situation.</i><p>Paul, I think, is talking about achieving outsized outcomes in relatively shorter timeframes (as the timing is just <i>right</i> to be investing in learning this tech) for high agency folks who can <i>also</i> afford the ordeal in wanting to maximize for impact & ambition. Of course, there's real risk one may get no where, but even in failure, given you were building the LLM yourself, you might end up with other adjacent, high reward opportunities.
> The reality is that an incredibly small minority of companies in the world do any real training or optimisation.<p>At the scale you are probably imagining, this is true - but take the hype out of the OP and what you have is just someone saying the field of data science exists and is growing.
1000%
Did you read his comments on this? It's not to actually do LLM research stuff, it's to trigger and unlock ideas.
I think his point is, if LLMs are the future (like computing is the future in the 90s), you should be an LLM-expert (equivalent of becoming a software developer).<p>I can see the point. It's unlikely that a 2.4T LLM will be integrated into, say, a pesticide drone. You'll still need some kind of LLMs to achieve maneuvers that "normal" programmings can't achieve.<p>But what if <i>everything</i> basically turn into that? Essentially, instead of build me a web app to solve X and do Y, build me an LLM to serve X and do Y. (unless the current LLMs are able to do it end-to-end but then they can hardly write coherent software/personal opinion).
This is like saying, “teens shouldn’t learn how to make their own game engine because no one is hiring for that.”<p>You’re missing the point. Understanding how Unity works fundamentally makes you a better Unity dev.
> You’re missing the point. Understanding how Unity works fundamentally makes you a better Unity dev.<p>Writing your own game engine makes you realize that the Unity engine is not really that well written....
Except knowing how LLMs work don’t actually provide much understanding for using them. People don’t use LLMs the way we’ve built on most other tons or platforms. It’s more learning Unity hoping to be a better gamer.
That's not true, because everyone, everyone, everyone seems to want to do training. Which results in a 50 person company training, say, a voice model that then fails, because it's just not good enough.<p>In reality the problem is that it gets blasted out of the water by a much worse architecture trained on 10000x the infrastructure. And while I'm sure the freshly brought in ML student came up with a 10%, even 30% better architecture, it just doesn't matter. (and never mind that even OpenAI hasn't really solved a voice model yet. Try it. It can probably match 2026-quality call centers, but it's no substitute for an actually empowered human)<p>... and yet, if you look at what hyperscalers are getting paid for ... comfortably more than half the income is training. Which makes no sense on so many levels.<p>e.g. <a href="https://valueaddvc.com/blog/inference-chips-vs-training-chips-why-the-next-semiconductor-race-is-different" rel="nofollow">https://valueaddvc.com/blog/inference-chips-vs-training-chip...</a> (I get it, not great first source, but st
The big question is whether companies hold enough proprietary data to do useful things that for e.g. Anthropic, etc. can't easily replicate.<p>For some very niche cases I think this is probably the case but for the vast majority, the company's data isn't as useful as they think it is or anywhere near the size needed.
Everyone says they want to do training, because it's sexy and an easy way to justify raising mad funding rounds. Some manage, most don't.<p>I don't know where you are located, but in EU, in China, and yes even in Silicon Valley, the vast majority of companies do not do any real AI engineering. There's nothing wrong with it, it's just not a smart path for most purposes. You can do amazing things without training, and if you try to train, you cannot get anything amazing unless you burn millions.<p>Very few people can afford to play the long game and cross that dessert. And, sure, you will not get far without good engineering, but good engineering is definitely not sufficient and is not the primary bottleneck.
[flagged]
[dead]
[dead]
[dead]
I'm (more than) twice that age, but I've spent time learning this exactly this from videos by Andrej Karparthy and from books by Sebastian Raschka.<p>I didn't do it because it was useful to me in a practical sense. It's because LLMs are fascinating and I want to know how they work. From that perspective it's been a great experience. I have afirm grasp of the basics. This makes it much easier to understand frontier concepts like compressed latent attention. I can follow the field and understand it.<p>Not sure I would have got as much out of it at seventeen. I have a lot of background and experience which made it much easier to learn. I wasn't struggling with the linear algebra or with python. I already knew pytorch and neural networks. That helped a lot and I covered these tutorials fast and could skip over large sections. A few evenings and the odd weekend day over a couple of months was enough for me.<p>For seventeen year olds the tutorials are good enough to make it possible to learn this but it would have taken a lot longer to understand. On the other hand I would have learned a lot more. I think I would have learned a lot of valuable stuff.<p>However I also think 17 year old me was studying for his A levels and probably this was right choice in terms of maximising future opportunities. I'm not sure I think learning about LLMs instead is sensible. Indeed it might be bad advice. But I can absolutely agree with the sentiment.I think 17 year old me would have wanted to do this too.
While knowledge is always great, I would encourage people not to seek advice from successful people like this (survivorship bias).<p>Moreover I am not sure it is even good advice? Would you advise a 17 y.o. to learn how transistors work or how to code (i.e. is LLM training the right level in the stack)? LLM training, a discipline where relevant work is already out of reach for 99.999% of budgets really as essential as this post implies?
> I would encourage people not to seek advice from successful people like this (survivorship bias).<p>Personally I don't see the problem, as long as you're aware there is survivorship bias involved here.<p>What's the alternative really, seek advice from unsuccessful people? That seems worse :)<p>Personally I do both, read about what worked for people, also read about what didn't work for people, then ignore both and do whatever the fuck I want.
It's best not to take advice on direction of careers from anyone. It's better to find and work on things that interest you, and then take advice from people that are amazing in that specific field. In mid 2000s in Australia all the "top people" were telling me not to get into a software engineering career because it was dead. It's certainly challenged right now, but it took off during those 15+ years.
“It's best not to take advice on direction of careers from anyone. It's better to find and work on things that interest you, and then take advice from people that are amazing in that specific field.”<p>What bucket should I put this advice in?
Understandable irony. The way I see it, there are no buckets, only context surrounding the advice and your awareness thereof. If the context is unknown, best to take it with a big chunk of salt... Or spend time familiarizing yourself with the context if you think it's worth the time/effort; working on things that interest you is a nice way of doing that.<p>But some advice is less specific than other advice though. Some things are always stupid, and some things are always smart, if you look at the context of our world and society. I find myself pursuing these "fundamental truths" with great interest lately, especially now that the world is changing so quickly.
A lot of the advice that is relevant for top performers in a field has no relevance to an average person interested in that field.
> It's best not to take advice on direction of careers from anyone. It's better to find and work on things that interest you,<p>Agreed, my previous stated "ignore both and do whatever the fuck I want" approach has worked out very well for me in life, people should probably focus on identifying better what their gut tells them, rather than what randoms on the internet thinks and writes.
>What's the alternative really, seek advice from unsuccessful people?<p>Seek advice from the averagely successful people, since that is statistically what you're most likely to be.
<i>> seek advice from unsuccessful people</i><p>Intuitively, I would guess that they have a better grasp of what made them fail than successful people have of what made them succeed.
I don't know; I think people in general are just not great at this. Successful people tend to underrate luck and overrate the brilliance of their own decisions, but the rest of us are prone to either reversing that and blaming everyone but ourselves, or being so determined to take accountability (or just depressed) that we become overly self-critical, or simply not understanding why things happened the way they did and reaching for any explanation that resolves the chaos into something narratively satisfying.
Intuitively, that'd make sense if that unsuccessful person eventually found success, otherwise who knows if they actually picked up what made them unsuccessful in the first place?
I think it's reasonably common for people to understand what makes them unsuccessful but be unable or unwilling to resolve the problem.<p>For me, I know that one major reason for putting an upper limit on my success is my inability to form effective professional relationships. I know I should go to events and talk to people and use these relationships to my advantage, but that's just not something I've ever been good at, and I find it so incredibly unpleasant that I also just don't want to do it.
Yep, many successful people greatly underappreciate the effect of simple dumb luck in their lives. And often they just make up complex reasoning chains, even wholly believing them, that more of it was in their control/talent, etc.<p>Nonetheless, there are many successful people I would gladly listen to for advice, though they are often successful in a different meaning than what venture capitalists would use (e.g. parents with great kids, managing to keep a healthy work-life balance, happiness, and maybe even having time to spend on some cool hobby project -- you are heros!)
same do unsuccesful ones. They usually overappreciate luck of succesful people from what I see, as a way of coping, many say "that guy was just lucky, i was unlucky". If you dig deeper, that person wasn't just "unlucky", they lacked the persistence, work ethics and other qualities compared to other succesful ones.<p>Of course, there are degrees and exceptions on every side, but the coping mechanism is very strong among people.
If there's a single unifying fault to the unsuccessful people I know (for any definition of unsuccesful, including my own failures) is that they're precisely bad at working out why they failed - if they even think to ask the question. The successful people are generally much clearer that it was studying hard or networking or natural talent that contributed to their success. The latter group may have people who don't realise luck played a part, but they massively increased their chances of good luck by the aforementioned behaviours.
Luck is a huge factor in success, so that logic leads nowhere. If one of the quacks selling how-to-succeed recipes had got it right, everyone would do it by now.
Kids don’t even know what it is, even less actually feel what it is. At 17 you think that it won’t hit you, that you will be the one to survive until you don’t.
> What's the alternative really, seek advice from unsuccessful people?<p>Do both. Get advice from successful and unsuccessful people and take the diff
> What's the alternative really, seek advice from unsuccessful people? That seems worse :)<p>Learning from their mistakes (where it did go wrong) can be as valuable as.
The better alternative is to always ignore unsolicited advice.
I would seek advice from people who have a theory of why or why not they were successful.
A lot of those results were happening in very specific contexts and usually should not be regarded as a blueprint, but as inspiration to whatever I do.
The unsuccessful person is much more likely to be actually connected to the daily needs and struggles of a normal person. Paul Graham hasn't seen the inside of a grocery store in the past twenty years.<p>I know who the 17 years old is closest to.
>Personally I don't see the problem, as long as you're aware there is survivorship bias involved here.<p>Pretty tall ask, especially if this is targeted at 17 year olds...
> Personally I do both<p>That’s what you get from listening to “successful people”. You get to learn about all the things they tried that failed, then the things that did work on that 24th try, which was successful.<p>The “survivorship bias” people always seem to assume that the “survivor” lucked into his fortune on his first try ever, so he can’t have learned anything, so we don’t have to listen to him. But that’s seldom the case.<p>I’ve written about this before:<p><a href="https://expatsoftware.com/Articles/survivorship-bias.html" rel="nofollow">https://expatsoftware.com/Articles/survivorship-bias.html</a>
very good read. I get into the same discussion sometimes with people. If you're successful you're automatically labelled as "lucky".<p>There are successful people that failed many times before being successful. Somehow these days, previous failures and persistence seems to be ignored and they focus just on the luck you got on the 20th try.
> seek advice from unsuccessful people? That seems worse :)<p>Not sure why would you think so.<p>Inverse reasoning is very powerful, and unsuccessful people can give you plenty of "don't do this mistake", which the survivors would not even think about.
> unsuccessful people can tell you teach you plenty of "don't do this mistake",<p>But how can I know for sure that that particular mistake is actually why they were unsuccessful? Has exactly the same issue as listening only to successful people as they hardly know what actually made them successful most of the time, but they still compose large blog posts with their reasoning for why.<p>Again, I still think my approach of reading both but then regardless go my own way is the preferable approach, at least for me, ymmv.
Seek advice from people who have had a normal level of success. Not a one-in-a-million level.
Veritasium has a good piece for this kind of bias:<p><a href="https://www.youtube.com/watch?v=3LopI4YeC4I" rel="nofollow">https://www.youtube.com/watch?v=3LopI4YeC4I</a><p>An advantage that is not "advisable", like being born in january, in a rich country, in an above average family, or just having luck, might have more influence on the outcome than any conscious action. It is almost sure that one-in-a-million level people only edge over the other 999,999 they competed with is just "have more luck".
As someone who teaches AI at a university, understanding the basics of how the LLM works is really invaluable to understanding where and how they'll be applied.<p>You're right that it's probably a little too in the weeds, but it's also a nice clear and fun objective that teaches you the basics. Like building a TODO list in JavaScript to learn webdev or a Gameboy emulator in C++ to learn how a CPU works.
I would encourage young people to be born rich. It's the best time in 100 years to be advantaged. Why waste your potential by having your labor stolen?
> Moreover I am not sure it is even good advice?<p>I think it is. He isn't saying to learn how to train a LLM so that you can go on to train LLMs. He's saying to learn it so that you gain a deep understanding of how LLMs work. Ordinary startups can still benefit from things like training or fine tuning highly specialised smaller models, knowing how to select and configure an appropriate model for the task at hand, knowing what software to use and why, understanding what's going on behind the scenes instead of treating everything like a black box, having a higher level of intuition about LLMs generally, etc.<p>Most computer science courses do in fact teach things which are lower level than coding, such as how transistors work.
IMO these kind of advices never matters. Any individual still needs to make tens or hundreds of little decisions (every day) themselves, and that's what really makes all the difference.
It cannot be bad advice because learning something even slightly valuable is always good in a vacuum (as you mentioned)<p>Since he has not given any reference point it is equally as good advice as „just learn everything slightly valuable“.<p>So the real question is: „What should I not learn in favor of learning this.“<p>Or in other words: His comment is not usable because it can mean anything or nothing.
> Would you advise a 17 y.o. to learn how transistors work or how to code<p>how many of us out here are doing work directly in what we got a degree in? I majored in economics and now I'm a CTO.<p>I would absolutely advise a 17 yo to learn how to code, understand how transitors work and how to code an llm. even if he never works on llms, you basically end up with a kid with applied knowlege of statistics, math, physics hardware, logic and a whole lot of practice in critical thinking.
As a 17 y.o (way back in the last century) I didn’t need to be advised to learn about transistors. I just had a thirst for the knowledge.
I would encourage everyone to learn something about transistors. They are one of mankind’s most useful discoveries.
Right - like advising people to learn nuclear power back in 1992. Perhaps a good idea, but super specific already. My gut is ML/AI is even more complex in 2026… I can’t even remember all the abbreviations and the new ones emerging. And every sub component, such as attention or embeddings, are actually a discipline of its own already.
I just don't think he realizes how saturated it got over the years. Or maybe he knows at a conscious level, but not subconsciously.<p>In 2000 (his era), it would have been really smart to study the source of Linux or Apache. Would have paid dividends over decades. Cuz that knowledge was so rare. The number of people hacking on LLMs now dwarfs the number of people hacking on web servers 30 years ago, by several orders of magnitude.<p>And if you turn back the clock even more, I mean just even having access to a computer, let alone owning one, would have put you at a massive advantage.<p>I don't know what to call it. The pioneers should be respected obviously, but at the same time you need to understand that for them, the game wasn't <i>nearly</i> as played out as it is now.<p>I just don't think you can afford to be dicking around with LLMs like you could afford to dick around with random Linux distros 20 years ago. Too many people willing to do it for free these days.<p>You don't wanna end up being the 2030 equivalent of a certain SNES emulator developer, or maintainer of a package manager for jailbroken iPhones, I mean the list goes on and on. Being a hacker doesn't automatically give you a path to being rich, or even making a decent living. It hasn't been that way for a while.
>Would you advise a 17 y.o. to learn how transistors work..<p>Hell yeah. Transistors are pretty awesome.
> Someone asked what I'd do if I were 17. I'd learn how to build LLMs from scratch, and then train ones as powerful as I could with whatever hardware I could get access to.<p>He's seeing a future for models running on everyday hardware just capable enough to do what the use case requires.
I agree more with Yann LeCunn's salty reply. Over long run, knowing how autoregressive language models work from scratch will be just one step in having foundational understanding, and they might become dated... like knowing how a CRT monitor work. Something of historical interest and good for learning, but not crucial to being well-rounded.<p>There are other types of models like diffusion models right now that are showing more efficiency and have a higher ceiling for improvement. Understanding math and fundamentals are more important.
When I was 17, I wrote poetry and learned how to play guitar. I’m really happy that I did. That’s what I would do now, were I 17. I feel slightly bad for people who didn’t.
I'll use this post as a shameless opportunity to tell more people about a little side project, I made:<p><a href="http://languagemodelbuilder.com" rel="nofollow">http://languagemodelbuilder.com</a> teaches you (in a few hours to days) how to build an LLM from scratch. It's entirely free, without accounts, and without data collection.
Beautiful work! What prompted you to build this? Pun not intended.
Nice! Please post screenshots and videos for people not ready or able to install yet.
This is the coolest link I saw on HN the last 2 months.
You rock man !<p>Thanks a ton for building this.
This is great, thank you!
I wish this was available for young folk in Romania: As it happens: macOS: 2.70% of desktop operating-system usage in Romania; OS X: 2.55%; Combined Apple desktop share: approximately 5.25%; Windows: 90.65%; Linux: 4.02% or so AI tells me.
I am kind of amazed how negative the comments are here, especially on HN.<p>Learning to hack something together in high school using the latest technology (vacuum tubes, radios, microprocessors, web/javascript) has been a common theme in the tech world for generations. With LLMs and online tutorials, this isn't even a difficult suggestion. Do people think learning new tech is somehow wasted effort?
If he has said "to learn the math and programming skills needed to understand how to build LLMs" it'd have been <i>much</i> more positively received.
> I am kind of amazed how negative the comments are here, especially on HN.<p>I can't recall or point out exactly when, but there is a stark before/after moment where the opinions of anything pg went from "Interesting and maybe true in some ways" to what we see today, lots of knee-jerk reactions and hardly any comments about the actual content.<p>Hazarding a guess, I think the moment Altman became the CEO and later during COVID, the sentiment seemed to have been shifting towards what we see today. But this is all based on hazy memory, rather than looking at the data. I'm sure there is a blog post waiting to be written about analyzing the sentiment of comments to PGs articles on HN, and you'll see a shift somewhere.
Imo it's a breakdown of trust of the startup ecosystem as a whole. Repeatedly startups have enshittified and it's become undeniable that the investment apparatus around startups is partly responsible. We have seen a great driver of uncreative destruction, industries undermined, small businesses undermined just to drive masses of money into few pockets - less fairness for the people working in what is now the gig economy and ultimately prices and other costs that end up as high or higher than they were before for consumers. Not to mention the whole AI/OpenAI situation which many perceive as threatening their skillset per se, essentially tearing up the social contract that existed on this site.<p>The sycophancy on HN is starting to break down because there is a higher proportion of users sceptical towards the outputs of the VC and wider investment world than ones who believe they're potential beneficiaries of it.<p>Tech industry people are becoming less interested in HN as a warm handshake into the startup world because, frequently, they're disgusted by it. And this reflects on the sentiments people post on PG's articles.<p>Increasingly if those at YC want the same kind of low-bar praise they got before, they will need to get it from machines.
Because at some point in life everyone gets tired of fairytales. He started mending the anecdotes to his content which always rubs people the wrong way.
So, this comment of yours obviously isn't in the "knee-jerk reaction" category of comments, I suppose? What exactly from the linked tweet(s) are fairytales here? There is hardly any text at all, so strikes me as a comment about previous pg content, but then this would be one of those comments I talk about? Very confusing.
Your hand waving doesn't make it knee-jerk. It's just what happened to his writing since COVID. He goes for more of a shock and awe style and not everybody likes it. He's been writing for over 20 years now, hasn't he? His style has clearly changed, and an changing style attracts a different audience so it's no surprise his original readers might not connect with his newer work...
> I can't recall or point out exactly when, but there is a stark before/after moment where the opinions of anything pg went from "Interesting and maybe true in some ways" to what we see today, lots of knee-jerk reactions and hardly any comments about the actual content.<p>Hard disagree. This submission is still being highly upvoted, while another recent post[1] on the harms caused by Graham’s fellows[2], with a fairly tame comment section, has been flagged. That is a constant on HN. It’s not a fluke, it’s as predictable as the sunrise and getting more pronounced.<p>I’m sure we’re both biased in our perceptions. Mine is that HN in general (certainly more than any other website) used to worship[3] everything he wrote, together with others like Musk, until things started to really go to shit and many eyes have been opened to the effects of the unfettered greed of rich tech guys out of touch with reality.[4]<p>[1]: <a href="https://news.ycombinator.com/item?id=49411762">https://news.ycombinator.com/item?id=49411762</a><p>[2]: A better English word is escaping me.<p>[3]: That word I choose hyperbolically but deliberately. It definitely was not “interesting and maybe true in some ways”, it was much more hardcore than that.<p>[4]: That is not “knee-jerk” but a slow realisation still ongoing.
> Mine is that HN in general (certainly more than any other website) used to worship[3] everything he wrote [...] It definitely was not “interesting and maybe true in some ways”, it was much more hardcore than that.<p>I guess it depends on what submission you look at, previous comment of mine solely based on memory. Now I went to <a href="https://news.ycombinator.com/from?site=twitter.com/paulg">https://news.ycombinator.com/from?site=twitter.com/paulg</a>, clicked "More" a bunch of times, and seems my memory was more or less correct, none of the submissions I clicked on are "pg worship" (hyperbolic or not). Just one example: <a href="https://news.ycombinator.com/item?id=19418701">https://news.ycombinator.com/item?id=19418701</a><p>Maybe you need to enable "Show Dead" or something? Pgs articles on HN definitely never was free of any critique in the HN comments, just like any article. Although I do agree with you that it used to be different than it is today, and same with Musk too, and Altman, and probably more individuals, where they were lauded before but now pretty much just mentioning them poisons the conversation.
> Just one example<p>That post has barely any points and comments. It’s not a good indicator of general sentiment, it’s just an indicator of people who were on HN at that time.<p>> Maybe you need to enable "Show Dead" or something?<p>I have it enabled.<p>> Pgs articles on HN definitely never was free of any critique in the HN comments, just like any article.<p>Of course. I <i>very explicitly</i> wrote “in general (certainly more than any other website)”. That does not mean “always”, or “never”. HN is not a hive mind, there’s never going to be 100% agreement. The <i>general trend</i> is what’s being discussed, and we both agree that in general the sentiment on Graham used to be higher. We’re just disagreeing (we may be able to find ourselves agreeing through tough thorough thought, though[1]) on where exactly it is now and how to interpret it.<p>[1]: Sorry, can’t believe there was a real organic opportunity to use that sentence, had to take it.
Indeed, no hard disagree, merely details :) Overall you're right though, general/overall tone definitely shifted hard for a bunch of individuals over the years.
> Indeed, no hard disagree, merely details :)<p>Hard agree! Though the discussion was short, I thank you for it. Good start of the week, I wish you a good one.
I think it is more about people are a bit sick of filthy rich people giving this kind of advice. I would also not read anything he preaches.
His writing was overrated: It's that simple. People now see his blog posts for what they are: decent blog posts.
Completely agreed. The point is the knowledge, the learning and the journey. If a kid has a passion for building or toying with LLMs, then of course, by all means, please start tearing them apart or even build and train your own model. You'll learn a ton, even if you won't necessarily end up using it here and now. The learning experience will compound and of course that will be useful.<p>The above is, after all, the whole genesis of the word 'hacker'. We should celebrate that.
Now I'm curious, do people actually tried to hack vacuum tubes or other big servers that's barely 1MB RAM? It seems like another thing that needs big investment to work properly, unlike those other techs where results can be shown even with little materials.
Yeah, there are retrocomputing hobbyists who mess around with sometimes-physically-large computers that were important many decades ago. I don't know if anyone is hacking on vacuum tubes of the kind that you could in principle build a computer with - there's a reason they became obsolete for digital computation almost as soon as the transistor was invented. On the other hand, I personally think it would be neat to try to build a CRT in a garage, which is of course a type of vacuum tube. I don't think this would be an <i>easy</i> garage project, but it does seem like might be tractable for someone who understand physical manufacturing and electronics well, has access to glassblowing equipment, etc.
> Now I'm curious, do people actually tried to hack vacuum tubes or other big servers that's barely 1MB RAM?<p>"Barely?"<p>It's insane to lump vacuum tubes together with servers with 1 MB RAM. My first PC, which I used for a decade, had 640KB RAM. And <i>that</i> was an upgrade from 512 KB RAM. My other PC had only 128KB RAM. None of these were considered the equivalent of (by then long dead) vacuum tubes in their day.<p>You can get a lot done in 1 MB.
> With LLMs and online tutorials, this isn't even a difficult suggestion.<p>Don't many of the commercial ones prevent you from using them to build LLMs?<p>I would say the reason for the negativity is not because it's a bad idea for a project, or that doing projects in general is a bad idea (it's not!), it's because it's a very specific thing that is not for everyone. The best thing about computing is the low barriers to entry. You can basically work on anything that takes your fancy. So those who are interested in ML will be drawn to learn about LLMs. They don't need anyone to tell them to do it. Telling everyone to do it reminds me of the "just learn to code" stuff of a decade ago. No, please don't, please find something you enjoy.
> Do people think learning new tech is somehow wasted effort?<p>No. But funnily enough that is a promise by some of the AI cretins and their boosters. Oh yeah best case scenario you learn how to build LLMs for us. We’ll employ you. And then ultimately that just becomes training data for the LLMs to do it themselves.<p><i>But why are people cynical?</i> they ask.
So I am 15. Is it worth trying to build my own archive (s-1.site) of strategy ideas? Do I have any real differentiation? Or am I just wasting time? I figure that I can use it as proof that I have some know-how?
I appreciate the sentiment and I'm pretty curious how I could train an LLM, even are really basic one from 5yrs ago without nuts hardware.<p>That said, I a good starting point for a 17yo is reading about perceptrons[0], then the basics of neural networks[1] (ex. 3-layer perceptron) then writing a program to train a 3-layer perceptron and classifying the MNIST dataset[2] - a dataset of characters.<p>This can anywhere between a day and a week and you will demystify the basics of neural networks and work your way forward with more advanced contemporary concepts.<p>Fun fact: any multi-layer perceptron neural net can be reduced to a 3-layer perceptron network.<p>[0] <a href="https://en.wikipedia.org/wiki/Perceptron" rel="nofollow">https://en.wikipedia.org/wiki/Perceptron</a>
[1] <a href="http://geeksforgeeks.org/deep-learning/neural-networks-a-beginners-guide/" rel="nofollow">http://geeksforgeeks.org/deep-learning/neural-networks-a-beg...</a>
[2] <a href="https://www.kaggle.com/datasets/hojjatk/mnist-dataset/data" rel="nofollow">https://www.kaggle.com/datasets/hojjatk/mnist-dataset/data</a>
Horrible advice. This may have been good advice 10 years ago, but not today. There are no positions for people who "kind of understand how toy LLMs work" because so many engineers do these days. Most of the real LLM optimization work is at the edge of research and highly proprietary and not something you could ever do without infra that costs millions.<p>But of course, 10 years ago this wasn't obvious.
you don't reach cutting edge immediately. you start with the basics
What would be better advice for a 17 year old?
Throw the computer and the smartphone out of the window.<p>Or more reasonably, the same old thing : use Linux, hack a little, why not learn programming basics. But learn to own your technology, fight against centralization of technology. The same old RMS story.<p>IDK where the tech industry is going, if there will be jobs anymore or not, but what I'm sure (and what have been the case for the last 10-15 years anyway) is that for most tech jobs, having good technical knowledge beyond the basics is pretty useless and will probably not be recognized.<p>If you can, stay a computer geek if that's your thing, but don't make it your career choice, the Eldorado is behind us.
Stay in school, get good grades.
Stay away from the internet and learn how to build a local business instead.
"it doesn't matter what you choose to study because no job is safe from being outsourced overseas, handed over to an indentured servant, or made redundant by a machine. you will spend your whole life surviving while the cannibalistic pedophiles who own everything invent new ways to make you own nothing. be frugal, don't get married, don't have children, do everything you can to stay healthy and independent, and you just might live a reasonably comfortable life."
Most reasonable advice I've seen under this post. Having graduated in 2025 I've experienced the horror of competing with infinitely many third worlders in my small country, who will gladly take 1/3 my pay and be serfs.<p>I really hope we get another hiring boom like in 2020 when I decided to study CS. Otherwise my career will be very rough. I love it and can't imagine doing anything else.
> don't get married, don't have children<p>Bad advice. You’re going to need family connections and loyal people you can bet your life on to survive if the rest of what you’re saying is remotely true.
Learn math
Yeah he's a moron with a lot of money, that's about it. I'm sure he's said the same thing about various other bags he had bets on throughout the years.
Don't really understand why people are saying this is terrible advice. If you're 17, you should be learning <i>everything</i>. At that age, your brain is a sponge and your energy levels are the highest they will ever be.<p>None of us know what the future of work, education, or AI is going to look like. But your best bet is to become a life-long learner. Be it LLMs, musical instruments, physics, or business.
I think a problem a lot of people are grappling with here is that due to LLMs and AI generally, it’s basically impossible to predict what the future will look like or what jobs will still be around.<p>I’d probably say something like: do something you enjoy and seems like it might be useful, but accept that the pace of change may mean that whatever you study ends up being irrelevant.<p>Whatever solution there ends up being to this, it’s not going to be one that an individual 17 year old can implement. We’re past the point where individual good and bad choices matter that much to economic outcomes.
I would advise any somewhat ambitious 17 year old to avoid tech and get into healthcare if they can stomach human interactions and bodily fluids. Sure, it is not all sunshine and rainbows, but there will still be plenty of work helping people who are ill or elderly. Even in the worst-case economic scenario, medicine will be a more socially rewarding and stable life path.
This seems like a much more interesting question to me.<p>Telling other people's children what to do is easy and basically doesn't have any downside to being wrong. With your own children things are a bit different.<p>So: what are people here with school age children telling their own kids about the future? If their kids ask, what kind of careers would they encourage them to pursue, assuming they have the skills and interest?<p>When I was last in the Bay Area, maybe about a decade ago the bookshops were full of titles like "Python for Preschoolers" (I exaggerate, but only slightly). Clearly at the time a lot of people working in tech thought that cultivating an interest in programming was going to be the path to being a successful (by some metric) adult. Is that still the case?
The real reason to recommend this is that they have an excellent chance of avoiding the "nerd-to-incel" pipeline that tech all-but-guarantees for its best nerds. Post GenAI boom, the social cost of working in tech combined with the coming collapse in high paying jobs, means that unironically coders should be learning a bit about coal mines.<p>In this regard, Healthcare is a polar opposite. It's pretty hard as a male nurse to not "accidentally" become a home wrecker.
As someone who's at a similar age and was interested in learning how to do this, there just aren't enough resources to do so. Most LLM research is in the form of academic papers, and there isn't any 'popular' way to learn these things, and besides that all said research assumes you have a B200 cluster ready to go. If you have weaker hardware (say an 8GB nVidia GPU, which is what I have) you're going to be limited to fine tuning small models or torturing yourself working the GPU for days per iteration trying to run things like <a href="https://github.com/karpathy/nanochat" rel="nofollow">https://github.com/karpathy/nanochat</a>, which is hardly an educational experience. Renting cloud GPUs is expensive, and at this age the most I could muster up for experimentation is probably $100 or so, which only gets me 25 or so hours on a B200 which just isn't enough. So why would I bother myself with this if I'm already at a disadvantage because of not having access to the right hardware and when surely there are better ways to spend my time? I concluded the only way to learn and be competitive is by finding work at an AI lab somehow (not happening at 17), or studying ML at the right university.<p>Putting this in contrast with programming, I learned coding when I was 8, and it was incredibly stimulating to learn because you can quickly iterate and there were thousands of books and YouTube tutorials that dumb everything down and teach you fundamentals. All you needed was a $300 computer, and you can learn nearly anything you want, without being gatekept from this or that because you don't have enough vRAM / an sm_100 GPU.
I think learning how to build your own agent harness framework from scratch and really understand each part of it, what are the modern components that a good agent harness are using these days, is more valuable than learn how to build LLMs, but that depends on what you want to do with your career.
Ah nostalgia. At 17, I learned how to write programs in BASIC on a mainframe.
When I was 17, I made bad cartoons and was in a band. If I were 17, I'd spend more time learning music and design theory. I also learned PHP at this time, but that was low on the list, friendships came first.
Yeah, no way.<p>I'd move to the middle of nowhere and work multiple jobs on a farm and in construction. Learn how to grow food, and build things. Meet the farmer's daughter, and marry her. Then, buy my own land, grow my own food, and build my own things.
Slight problem with step “buy my own land”.
Your working 2 blue collar jobs and paying rent are incompatible with this plan.
I already do this. I live in a small village where there isn't even a wired Internet connection (wireless only). I work full-time with LLMs, on own product ideas and client projects (all LLM led).<p>I started investing in farms, have 50 pigs and 100+ chickens now. We are planning to grow to 100 pigs and 2000 chickens in a year. We will start growing Shiitake mushrooms in a few months too.
I get this is basically advice for young founders and entrepreneurs, but i would ignore that request and encourage 17 year olds to spend time trying to find a happy medium between work and life.<p>Being a super rich and an unhappy workaholic, or a super-impressive engineer who wakes up one day at 45 and realizes they regret wasting half their life (I ran into way too many of these) is a much worse fate than "not being rich from your startup" and working a relatively regular job while feeling fulfilled and happy by more than just work.<p>Especially in the US, which is uniquely bad at this and encourages people to work themselves to death, mental health and work life balance are much more valuable things for 17 year olds to focus on than finding good startup ideas.<p>In case you think i'm being a bit dramatic, let's look at the state of 17 year old mental health in the heart of Silicon Valley:<p>"The City of Palo Alto and the Palo Alto Unified School District approved a funded contract to place 24/7 human security guards and monitors at all four local Caltrain grade crossings, including the Churchill Avenue crossing directly adjacent to Palo Alto High School."<p>(in case it's not obvious, it's because of suicides by high school students)<p>The 17 year olds do not need advice on better startups, and this situation will never get better if we focus our advice on how to be better at work instead of how to be better at life.
This will require redirecting the conversations.
Thank you. They need human contact, not more "sit in a room alone and get stressed as fuck for little ROI" tech bullshit. Unless the kid has a genuine, self-motivated interest in learning these things (a great, positive thing that <i>should</i> be nurtured), they should file pg's advice under "ok boomer."
Mr. McGuire: I just want to say one word to you. Just one word.<p>Benjamin: Yes, sir.<p>Mr. McGuire: Are you listening?<p>Benjamin: Yes, I am.<p>Mr. McGuire: Plastics.<p>Benjamin: Exactly how do you mean?<p>Mr. McGuire: There's a great future in plastics. Think about it. Will you think about it?<p>---<p>I love this scene because it so perfectly captures what it's like to be young and given advice, however well-meaning, by an older generation living in a world that no longer exists for the young. And it's ambiguous and trite enough to be essentially useless even if the underlying idea isn't terrible.
A teenager in China actually did that and got a paper accepted at ICML.
The interview podcast is in Chinese, but you can ask AI to summarize it
<a href="https://www.xiaoyuzhoufm.com/episode/6a8472b95aeb2a5712e8de78" rel="nofollow">https://www.xiaoyuzhoufm.com/episode/6a8472b95aeb2a5712e8de7...</a>
I built from scratch a sparse text embedding model trained on a 13T token corpus. Not the same as an LLM because there’s no transformer in the mix, but still I learned a shit load of things in order to solve all sorts of problems that emerge when you try to access big datasets and daily update tables with billions of rows. But if it wasn’t for a specific use case that I tried to solve I don’t think that whatever knowledge I gained could be utilized in the market. Sparse models are a very small niche and most people I’ve come across with similar knowledge are in academic circles, not business related ones. So even if LLMs are all the rage these days I doubt the demand for people who know how to build them is that high. Someone who knows how to setup an open weight model and expose an API might be more valuable to a company these days.
> Notice that what I would not do is try to start a startup. Instead I'd build the foundation of knowledge to base a startup on later.<p>There's also this little-known concept called learning things for learning's sake and not always trying to capitalize on it.
If I were 17 again I'd prepare to go for volunteering overseas after high school for 1-2 years (plenty of free options in the EU where you might only need to cover the plane ticket). See the world, you learn a new language, help others and then think about what you want to do.
Most 17 year olds I know don’t want anything to do with AI and see the entire industry as an existential threat.<p>Nothing wrong with learning the theory and understanding the papers. Getting to that point you’ll have to get your fundamentals down. Might be an interesting exercise.<p>But as a future? I guess we’ll see. I suspect the next financial apocalypse will determine if there is one. Another AI Winter that may outlast all others so far.
I'm not sure what 17 year old me would have done with YouTube tutorials for everything under the sun available.<p>I'm much older and less wise now, but I still afforded myself the opportunity to follow karpathy's tutorials to build a LLM from scratch. Got to play with a few ideas. Seen similar ideas turn up in frontier model work, which is quite gratifying.<p>There are so many ideas to try.<p>Currently playing with autoencoders that takes A and B and produce latents A', B', and C'. Reconstruction of A is from A' and C', B is from B' and C'<p>The idea is if C' can be made to improve both outputs, it must store as much information as it can about what is common to both inputs.
Yeah sure, learning the internals of a technology that’s being actively developed will probably teach you things that remain useful for awhile even if your learning outcomes could end up being different than what he’s implying.<p>Whether that makes a good long-term investment is quite debatable.
IMO, LeCun’s take (at the end) is far more forward-looking since it aims to gain more insight into what could come next based on what we know about the current state of the art.<p>All that being said, the extent to which these people capitalize on our tendency to be blinded by the halo effect is incredible. A constant stream of bite-sized aphorisms…<p>This is especially at a different level for early startup figures who happened to be in the right place at the right time and usually did little more than digitizing mundane, traditional day-to-day processes. Yet they’re treated as geniuses and prophets, with people hanging on their every word as though everything they say contains some deeper wisdom. PG and the like often strike me as broken clocks and they’re still profiting from having been very early players in the game, who had good instincts for commercialization and capitalism.<p>————<p>LeCun’s reply:<p>> I would try to figure out why LLMs can write my essays but not clean my bedroom.
Then I would study topics in college and grad school that could help solve that problem.
I'll figure out a set of methods and architectures beyond LLMs that can quickly learn to perform physical tasks as efficiently as humans and animals.
That last item is also what I would if I were 30, 40, 50, or 66 years old
I thought pg was trying to live forever. Has he learned LLMs from scratch?
I really don't understand Paul's reasoning here. Does he predict more scarcity on the model-level? That layer seems to be almost a commodity now + training is damn expensive.<p>If you are really 17, my advice is to identify use-case for AI (ideally relevant for businesses) that work most of the time and find ways to make them work pretty much every time. AI reliability is the scarcity right now.
Oddly, I just remembered I did the nearest possible thing to this when I was 17... back in 1991.<p>On an Amiga, I took various public domain text documents from cover disks and counted the probability of the next word given the previous word. Then spat out random sequences of words from it and printed them out. It was called "Splurge". Basically a very very simple single layer statistical language model.<p>Some of the sentences were randomly not bad sentences, which seemed amazing at the time!<p>That kind of thing (and Core Wars and Tierra etc) <i>did</i> lead me to getting a job at an artificial life startup at the end of the decade. But that was in turn about 10/15 years too early (no GPUs).<p>There's some lesson from this about timing, but honestly I've gained the most as a person when I did something that was fun, ethical and gained an audience. A tricky combination.
I get it the sentiment behind the post…but it has some “let them eat cake” vibes though.
I think it's more about the scar tissue (a.k.a. intuition) when building LLMs from scratch than whether it's transferable to job search. Maybe the person will decide they don't like LLMs and not develop that into a career, or maybe they become a researcher in another field because that's a better way to solve problems inherent in LLMs.
<a href="https://deeplearningwithpython.io/" rel="nofollow">https://deeplearningwithpython.io/</a> is a decent (free) read for anyone looking to follow this advice.
This might be of use<p><a href="https://github.com/raiyanyahya/how-to-train-your-gpt" rel="nofollow">https://github.com/raiyanyahya/how-to-train-your-gpt</a>
I'd learn a trade in all seriousness.<p>(Edit: And learn how honest business works)
With the hindsight of experience, the remnants of my 18-year old energy go “woah, that’s cool!” at plenty of engineering feats… and my decades-older second brain goes “well d’oh, I could’ve just learned a trade to work on that!”<p>I think the last one was seeing a skilled electronics repairman do surgery on a CT machine controller.
Depends if you’re 17 with rich parents or not.
You can build the program, but to train it is another beast, billions of docs, images, videos, which a mere mortal doesn't have access to<p>Second in line, build your own agent, that's more in our ballpark, then customize it to your needs, both virtual and physical
I'm making an assumption here, but I think paulg is implying that "learning LLMs" today is like the equivalent of "learning computers" in the earlier days. We could even divide civilization into two eras: Before Transformers (BT) and After Transformers (AT).
When I was not 17 at the times of GPT2, I decided to not bother with learning how to build LLMs because it’s too expensive for an individual. This escalated quickly.
If i were 17, I'd try to distinguish who to take advice from, and would definitely learn that VCs have interest to spread a specific agenda in their message. Also, I would get drunk and have as much fun as could, as the misery of working under the treat of being replaced by AI, would simply kill any desire to live past 25.
i was trying to build llms from scratch at 17. failed miserably because i did not know linear algebra.
ended up in a different but adjacent field. when chatgtp got big suddenly there were so many people doing llms and i didnt want to compete like that. so much of what was happening was hype, and that really turned me off.
am back to building llms from scratch, but like, its a journey teaching myself all the theory on top of my job. i have decent fundamentals, but i need a better grasp of all the advancements in the field in the past 5 years before i would feel comfortable designing anything. baby steps, essentially. am working on better understanding all the layers of an ai while implementing a rag on my local model as an experiment.
17 year olds should learn whatever theyre interested in but need fundamentals in order to do anything advanced.
I learned HTML when I was 17 in about 1995 and it's certainly taken me on a pretty fun career path. Less technical than LLMs for sure, but 'figure out where the industry is going and move what you're learning to there' is solid advice.
This essentially comes down to choosing the search for substance (here in the form of technical depth) over short term gratification and quick wins in life.<p>I think this kind of mindset should be taught way more in school so that people really appreciate learning a subject deeply.
The only realistic approach is to train on a limited data set which is probably less usable than the comibnation of a custom RAG + one of the many available LLMs.<p>Also many people/kids don't have access to proper "productive" systems anymore, since the whole computing and electronics industry shifted to make "consumer"-devices like smartphones or laptops made for netflix, gaming and spotify.<p>Breaking the barrier to build a custom system, install linux (or developer tools for Windows, MacOS) is already a complex AND costly task. It was just way simpler in the late 90s and 00s to get something working.
Writing, supervising and training LLMs are now the purview of... even larger LLMs. Optimising CUDA kernels; hand-writing SIMD assembly to speed up data loading; tinkering with your particular brand of DRAM to see if there's anything to gain from optimising for its memory topology and NUMA --- these are now the job of AI.<p>There is very little reason for humans to get all too engrossed in this type of work now, today, with the hope of being good enough at it to command a high salary in 3-5 years. AI can already do it incredibly well, and they can do it persistently and doggedly 24 hours a day.
telling a 17 year old to get into tech right now is horrible advice, literally telling them to get at the back of a line with a better part of a million more experienced people in it.
Telling them <i>not</i> to get into tech is also terribly reactionary advice. The truth of the matter is that we don’t yet know whether tech roles will be eliminated or if they’re just going to follow previous innovation breakthroughs where “one person producing way more work” makes software even more of a desirable industry to be involved in.<p>There really isn’t a very strong correlation between tech industry hiring strength and AI as of yet. Various studies that are out there haven’t even witnessed AI workflows contributing more than modest gains in software engineering efficiency. I.e., being able to write code 20-40% faster isn’t a seismic shift in the industry where everyone is getting laid off tomorrow and we’re all replaced by software.<p>Even with the questions surrounding the current job market, it’s still an incredibly good ROI career compared to so many other jobs out there.<p>For example, in my local area you can get a job as a registered nurse working nights in the emergency room and only make ~$115k.<p>I make almost double that telling an LLM what to do from my house in my pajamas during the day with less time spent in university.<p>Even if tech roles lose half their salary to automation pressure it’s still a really good gig.
“only make $115k”<p>This right here is why nobody is shedding tears for the massive employment crisis in tech.<p>You make double that shilling ai slopware while they work nights saving lives.<p>I would tell a 17 year old that the world will always need nurses, same can’t be said for guys sitting in their pajamas burning tokens.<p>What is happening now in tech has been a long time coming, and it can’t happen fast enough.
I just want to make it clear that I’m not assigning some kind of moral superiority to my financial situation and the amount that the economy values my labor per hour.<p>I don’t make the rules for how much each profession is able to make in compensation.<p>The world will always need nurses, but that doesn’t mean that it’s a fantastic career to get into if you have a neutral career preference and your primary consideration is university tuition ROI, expected compensation, work schedule, and day-to-day physical exertion.<p>My point isn’t to debate the virtues of each career, I am intending to stick to objective aspects of them.<p>In that sense, telling today’s kids that there’s no future in tech careers just because there’s a short term hiring slump is extremely premature. I certainly wouldn’t tell a kid who is passionate about tech to avoid the field just because the unemployment rate is currently 7%.<p><a href="https://www.investopedia.com/bachelor-s-degrees-with-the-best-and-worst-return-on-investment-12010130" rel="nofollow">https://www.investopedia.com/bachelor-s-degrees-with-the-bes...</a>
What else can he say? His business depends on that line continuing to exist.
Shaka, when the walls fell...
My kids CS teacher asked me what they should do after AP CS.<p>Previously they had a class where they'd build apps for other teachers. Like tracking when clubs are, etc. But now that's become easy for teachers to vibe code themselves.<p>I suggested they shouldn't prereq this class on CS. Heck invite anyone in interested in "building things" and they can get practice at building apps for other people / themselves. Maybe that becomes a gateway TO CS - people who want to learn how things work under the hood.<p>I suggested post CS class for the CS people should probably be building an LLM or something :)
A 17yo can train a small GPT this weekend. nanoGPT is a few hundred lines. Understanding why it works is the part that takes a decade.
What about doing abliteration, weight pruning, representation engineering, etc, directly to open LLMs instead?<p>Building an LLM from scratch has a hard split between a tutorial project you can complete in a weekend (that's useless for actual usage) and then a solid 1km high brick wall if you want to create anything actually useful from scratch.<p>Modified open models have a very active community around them, without the need to look much further than Hugging Face.
I'm usually a fan of pg, but this post is ignorant of modern AI technologies. Building an LLM from scratch is both a trivial and a useless exercise. There's probably in the range of 5000 github repos doing exactly that. What makes LLMs work is scale, and what makes engineering and training LLMs hard is also scale. And scale is not something you can achieve in your garage.<p>If the goal is to understand LLMs deeply, one would be better served by either joining one of the big AI companies or doing a PhD.
And to be honest, I think this journey should have been started 5 years ago, because right now there's too much competition.
Paul has the same problem just about every tech obsessed engineer has (including myself), he thinks everyone else loves computers too. They don't. At all. I''m a self-taught ex-bartender and I can't tell you how many grown adults in the service industry I've tried to get into computer stuff and they had 0 interest.<p>And younger people I meet don't even own laptops. I had a genz/millennial cusp friend who wrote all her college papers on her iphone.
I wonder at the “worlds” Mr Graham envisions, and what is their cardinality. Is this the only 17 yro reimagining, or is there an army of 17yro, of which this LLM curious persona is but one?
What would you do if you’re 30 years old now?<p>As a platform engineer being based mainly out of Australia/Hong Kong, opportunities seem to be getting less unless targeting high frequency trading or banking.<p>It seems like building a startup with the help of some AI tools might be the best bet.
The core point here is that AI is a massive thing (at the moment) so it's probably a good idea to understand it deeply. Not sure why people are so worked up about it.
Why is it important to train LLMs or even fine-tune them? LLMs have proven their point, costs are crashing, and there are more of them than most companies need.<p>The real value is to unlock meaningful insights and directions from existing data that is there inside companies.<p>I live far outside any tech city, so maybe I do not understand. But working with LLMs full-time, building for clients and tons of own experiments, I see no value in building on LLMs.
I would (and am) going into MLOps. Not just the general infrastructure/systems administration but how to do inference optimization, caching, quantization, memory pinning, vfio passthrough of gpus etc.
If I was 17 again I would sack off work and study and focus on chasing the opposite sex, without the angst I had at the time. I don't know a single person who regrets having had too much sex when they were young. I would not build an llm, too hard, too expensive to run. Learn how easy life is if your morals allow you to grift money of vcs into your own funds and retire.
Easier said than done - where do you get the B300s from?<p>Better to start working with harnesses, evals, statistical analysis, etc. - where you don't need the huge hardware for pre-training etc.
This might be a helpful resource: <a href="https://laurentiugabriel.github.io/token-town/" rel="nofollow">https://laurentiugabriel.github.io/token-town/</a>
I think pg answered the question as “what I’d do as a project” and not “what I’d do as a career.” So the critical comments are kind of missing the point, IMO.<p>I don’t see why learning how LLMs work is a bad project for a 17 year old.<p>Optimizing your entire career and the next decade+ of your life on LLMs? Yeah, probably not ideal. It’s almost always a bad idea to make long term decisions based on current trendy things.<p>And since everyone is using this topic to give their ideal advice to 17 year olds, my advice as a mid-30s guy: seriously consider becoming highly skilled at a specific thing, and don’t be scared off by the idea that it’ll take 5-10-15 years to get there.<p>When you’re 17-25, the timescale of a decade seems infinite. But it’s really not, and a decade spent “exploring and keeping your options open” sometimes just ends up with you being <i>pretty decent but not amazing</i> at a lot of random things.<p>Sometimes I wish I had just become a carpenter, chef, electrician, etc. – a specific skill set that leads to mastery over time, rather than the endless exciting-new-thing hamster wheel of working in tech.
If you're 17, and seriously curious about how modern day AI works, you might as well just sit down and look at a couple of courses on linear algebra + calculus, machine learning, deep learning, and more LLM specific deep learning. Those courses will teach you how to go from writing your first perceptron to a MVP language model. But also so much more.
As an aside, for someone interested and who's an absolute beginner, can someone please recommend good resources on how to build LLMs from scratch? Thank you in advance.
1. Build an LLM from Scratch by Sebastian Raschka (<a href="https://sebastianraschka.com/llms-from-scratch/" rel="nofollow">https://sebastianraschka.com/llms-from-scratch/</a>)<p>2. LLM from 0 to Hero, and nanoGPT by Andrej Karpathy
I second 1. I'm a newbie in neural networks and I think it's an excellent book! One of my barriers in ML is the resources, I find them overcomplicated or too simplistic without a mid term. It's not the case of this book, everything is well-explained. Neural Networks aren't fun for me, but this book makes it very interesting.
Stanford CS336 is up on youtube from Spring 2026.
Check the front page of this god forsaken website a few times a day and you'll get about 10 different posts a day about it.
... and it would be totally pointless.<p>I mean first that is already what plenty of 17yo are actually doing, because that is what they do at school or in parascholar activities. There are already countless of such tutorials where you can do that in an afternoon.<p>The pointless part though is precisely why Amazon and others are hunting for rare books, all the low hanging fruits have been picked already so just training a bigger model will simply mean burning more energy and money. Sure training a small one for the basic principle is a great pedagogical thing, training another one, medium, then maybe a large one, is also good in term of learning the process and architecture, but one should not expect it to be useful out of that context.<p>Pure players are precisely doing everything they can to corner the market by making their own scale unreachable by others. Smaller players with access to lesser infrastructure are thus betting on different market, e.g. embedded systems.<p>17yos should definitely build their (L)LMs from scratch and whatever bigger model they can train for free, or for cheap, but they should not expect that to bring them any riches.
Why would 17 year old do something that only brings them money? I do not think Mr Graham here is advocating for the path that makes most money as a result of learning how to train a model. I assume that tinkering and learning about LLMs is what enterprising 17 year olds will do to discover ways they can get a competitive edge or further the SotA with their insights further down the line.
No wonder the state of everything when 17 year olds are getting pressured to be “enterprising” and “get a competitive edge”. How about learning to be empathetic, respecting your fellow humans, caring for the place you live in, enhancing the lives of others? We shouldn’t be teaching 17 year olds to be greedy, selfish, self-aggrandising blowhards like Zuckerberg, Musk, and Graham. They are not good role models for the future of humanity.
he does mention that it would later on be about building a startup, which is about making money
It seems obvious to me that no matter what your age right now, you should want to be as broadly educated and intellectually curious as possible.<p>You should want to train a LLM from scratch as an intellectual curiosity itch that needs to be scratched.<p>The idea of learning one hot skill that has a pot of gold waiting at the end of it was a brief moment in time that came and went.<p>When I was 17, we would have said obviously support vector machines are the future. Neural networks overfit and don't work.
Paul G is not writing this for a general audience of your run of the mill “engineer” hoping to be employed by someone. He is writing it for future founders. What knowledge / skills you need to develop today to be well positioned to have a startup worthy insight when you are 24.
LLMs are the new compilers.<p>I don't think you can really call yourself a developer unless you at least have an idea how to build a more complex software project like a compiler, and maybe have built a toy one either at uni or for fun.<p>It's not clear how long this LLM age of AI will last (to be replaced by something better), but nowadays any developer should at least understand the basics of ANNs, and more than just the "hello world" of a cat vs dog CNN. An LLM/Transformer is maybe the equivalent of a compiler in that regard - something that we all use and is complex enough to present a bit of a challenge. You should at least understand the basics of how an LLM is built, and maybe building a toy LLM will/should become the new Comp. Sci. degree toy compiler replacement.
An LLM isn’t hard to make - the training data is hard to get and prepare.<p>The big companies stole the data. The average person can’t do that
It’s probably wise to learn how to build one to understand what you’re dealing with. However, if I were 17 I would lean how to apply an LLM to a problem instead of strictly building one.
If you’re 17, you might learn hands on knowledge, like tacit knowledge in areas such as lathes, precision engineering, metrology, and other very niche fields. You can also try climbing and explore arts like music, painting, and drawing. And, of course, spend some time in nature.
LLMs are incredibly boring to me as a technology..Not in terms of what it can do, but how it works.
Same for me. Neural networks in general.<p>I first studied them in 2011, and I was like what? Just a bunch of partial derivatives?<p>I keep looking at AI to check if now it's something else but it keeps being gradient descent.<p>Okay, it's great that you can perform miracles using gradient descent but that doesn't make it captivating in any way.
Can any1 share any resource to learn that skill. I want something that has been tried by you. I can too search on the internet...
certainly better than wasting time with harness and agent workflows that will become irrelevant at the next evolution, same thing happened with 'prompt engineering'
i am a small fan of pg, nevertheless i find this to be an exceptionally good take and it is strange to me to see so much piling on to this one in particular here.<p>learning about llms is not useful so that you can make llms later, you want to learn about it so that you can work on next generation architectures. llms before long i imagine will be left in the dust by ebm / physics oriented models especially that can have an embodied understanding of the world. but a lot of things you learn about them are transferable by doing something like this
Didn't his swiss watch essay say he'd essentially leave the industry because there will only be bloat from now on?
Why are people so negative about this? It feels like a fun project and at 17 the stakes are not really high. Something one could easily do on summer break in a couple of weeks.
If only we could train LLMs on commodity hardware, right?
As if you couldn't learn what LLMs are at any Age.<p>The core technoology is pretty basic, developing a rudimentary understanding for why the individual parts work as well as they do is tricky.
I read a bit about how LLM works, but as a hobbyist it is pretty frustrating that I won’t be building anything useful without throwing a lot of money at it
I'd just build an agent, it's easier than you think and you'd learn a lot about the "magic" of LLMs.
Seek advice from nice people that you actually know instead of rich people on the internet.
It’s really disheartening to see how many people don’t know shit about LLMs, by reading the comments… ironic given what OP is trying to say
LLMs will take away all jobs except yours, LLM maker. Keep at it, you're safe!
2 years before everyone was doing custom training. What happened to all those today when frontier models itself become more powerful than custom trained ones?
17 is such a fantastic age to be free and experience the world, you won't get that much of an advantage as these FOMO groomers are selling you into if you start now versus later.<p>If you are 17, go be yourself, whatever that is, in whatever way you want that to be, but do it so authentically and fully. Be unapologetic about what you love and what motivates you, and pursue that with passion and commitment.
Paul Graham:<p>>"Someone asked what I'd do if I were 17. I'd learn how to build LLMs from scratch..."<p>That's funny, Paul Graham, because if <i>I</i> were 17 again,<p><i>I'd learn how to program in LISP.</i><p><a href="https://www.paulgraham.com/rootsoflisp.html" rel="nofollow">https://www.paulgraham.com/rootsoflisp.html</a><p><a href="https://www.paulgraham.com/iflisp.html" rel="nofollow">https://www.paulgraham.com/iflisp.html</a><p><a href="https://www.paulgraham.com/hundred.html" rel="nofollow">https://www.paulgraham.com/hundred.html</a><p>(And/or other <i>LISP derived languages</i>... Clojure, Scheme, Racket, TinyScheme, etc.)<p>I guess <i>"the grass is always greener..."</i> as that old expression, that old "chestnut", goes... :-)
Would building an LLM from scratch imply writing code by hand?
What are the best resources to learn how to build LLMs from scratch for 17 year olds?<p>I have my opinion on this but I'd like to hear the HN opinion, I will just say one thing:<p>If you are starting with little knowledge, like a 17 year old would, letting an LLM explain it to you is a terrible idea.
Why do intelligent people still use X? Thanks for the xcancel.com link!
I've come to accept that some people, when faced with impressive technology, simply want to use it. They genuinely have no interest in understanding how it works. Lately, its even become fashionable to shame people for trying to understand ("you still read code? gross, you know AI can do that for you"..)<p>I will never understand this mentality, to let yourself be so dependent on something you don't understand at all is to live like a child. But it is very common. I doubt too many 17 year olds will bother even trying to understand what an LLM is, let alone build one from scratch.
When you see anthropic job postings that say the job may not be available in a year... it really doesn't inspire much confidence. I say if you're 17, learn to weld, solder, and fit pipe...
That's applicable only if you live in the US.<p>if i were 17, i'd do these things:<p>1) LC until mediums<p>2) Calculus & linear algebra even if i understood nothing i'd just stare at vectors and derivatives.
I would not waste my time with yesterday's fad. The next unicorn generator will be something else.
This is about as intelligent as say "If I were 17, I'd learn digital electroinics". You would waste your time. Sure, in theory it's useful, in reality it's not that useful.
All the negative comments here are really missing the point. At 17, you should be building your foundation. Kids that can make a custom CPU, or retrofit an old car with an electric motor, or screw around with nuclear energy... these are kids that are doing it just to see if they can. It's not about jobs, it's about curiosity and stretching limits and seeing what you can do, who you can be.
When I was 17 I was building Windows Phone apps, bad decision on my part.
I also did, and had written code for/on other mobile devices before, and did write a large number of even very ambitious software for other mobile devices later.<p>I do not see any past constructive experience as a waste of time.
Incredible counterexample, but oddly relatable.
I'd probably have achieved techbro 'post-economic' status earlier if I focused on Android dev instead of the shiny (and new at that time) Xamarin for Windows phones.
God I almost invested in Xamarin after Windows Phone got aborted, I did spend a lil time on UWP, but thank god Flutter came out not so long after that. After all these years I learned to stay away from Microsoft tech stack.
I think this is weird advice.<p>Learn how to make language models from scratch, yes. But learn how to use them, in the context of other machine learning tools, on very <i>small</i> hardware.<p>When the bubble bursts (and I still tend towards thinking it could burst rather than be deflated in a manageable way), the focus will be on uses of AI that are not like the hyperscalars' products.<p>People will still be interested in useful AI being added to small things — assistive technologies, home security, garden monitoring, their phones and smartwatches, robotics.<p>Instead of reductive, reactive make-an-anthropic-competitor advice like this, what about advising 17 year olds to focus on broad, integrated, helpful AI — or on going back through eighty years of history to look at AI projects that failed and reassess them?
Curious, what would you do if you are a 40 year old?
I remember when I caught PG on reddit arguing with some guy who'd said something mean about him. He didn't reveal who he was. But you could tell from his history - his first post was from before reddit opened to the public.<p>Good times.<p>I wonder if you can dig that out of the historical reddit database. I'd like to see that again. I love how everything is recorded now.
yeah i loved tech so started learning circuity and soldering. but it was a waste of time i made my living learning how to program web applications.
If I were 17, I'd be going to parties, music festivals, chasing girls, and enjoying my youth.<p>But sure, make the kids even more depressed by telling them they need to learn how to build an LLM so they can get a job working themselves to death to make Paul and friends rich.
This community is filled with people who were obsessed with microcomputers or (depending on which generation) websites during their teenage years. HN was created by and is managed by people of this type. No doubt we're a minority, and the type you describe is more common, but if you're implying there's something wrong with teenagers being intellectually curious about technology, I can't help but think of the way "nerds" used to get put down and shamed in the past.
Paul Graham wants you to work for him, not be him.
Really? because I feel like LLMs are already pretty much a commodity, not to say their won't be advances in LLMs but I don't see the models themselves being all that ripe for disruption the way that the web was and such. I'm guessing chip design and manufacturing processes will be more important than models in the future.
If I were 17, I’d learn how to invest and build financial literacy, and plot potential growth of my networth throughout my life, before even thinking about a career. Then smoke a bowl.
It seems this is poor advice in that it’s suggesting young people should focus on the current problem as opposed to future problems. Focus on the current problem can result in making some money but it will result in making the incumbents more money, which is not disruptive. Isn’t the goal of radical software startups to maximize disruption?<p>If the two current bottlenecks, for this LLM madness that could very well be a bubble, are processing capacity and accuracy (a second processing problem) then what comes next? Isn’t that where young people should be looking or are we just giving up on innovation?
I get the sense things have changed a bit since I graduated and there are lot more jobs in AI outside of academia these days, but it's still a very different field from other SWE pursuits, and it's not really accessible to hacker-minded people.<p>Learning AI isn't like learning HTML in the 90s then expecting to get a job at a tech company building websites. You can't just "learn how to build LLMs" and expect a frontier lab to hire you so I'd argue this is rather bad advise.<p>Additionally, unlike web development in the 90s you cant really do anything interesting yourself... All of the interesting/useful stuff will require huge amounts of compute and data so there isn't even much point in learning to start your own thing either.<p>As someone whose built many of NNs from scratch (hand written code, long before the days of LLMs), it's more or less useless knowledge if I wanted to work in a frontier lab or do anything interesting in the field.<p>I also think anyone thinking about going into a field which is basically a crossover of CompSci and Maths is absolutely insane right now. Even if you think there is a place for CompSci and Maths post LLMs, there's almost no chance anything you learn today will be relevant to the skills required in say 5-10 years.
It’s crazy how much survivorship bias gets repackaged as generic advice.<p>Wait no it’s not, that was always happening.<p>What’s crazy is that people still believe in it.
YC Combinator guy says that with a time machine he would learn to build the currently trillions-valued or whatever technology. Okay.
No, that isn't what he said.<p>He said, if he is 17 _now_, with no family, no commitments, no pressure to start a career, he'd invest his time to learn to build LLMs from scratch instead of trying to start a company.
I don't think individuals have the resources to build an interesting llm. The l stands for large. You need a dataset too. Llms are only interesting because theyre large<p>And it's basically a weekend project to put transformers together in a ML library and train it.<p>The follow up comment,train it to play a game also doesn't make sense? Llms Sony really play games and there are better ml approaches to do that?
On first glance, this is good advice and in general I'd give the same for this age group. With age, you'll see more of these scenarios come up and if you have experience and foresight, you can provide direction that may prove fruitful to young people. My son and his friend asked "what should we look into and learn?", about 15 years ago, I told them Python and Java; Python because of versatility and cryptocoins; Java for the long term stability in the job market. I despise Python personally, but I could see its potential and still do; especially for AI. At the end of the day, kids have more time than money and its great experience for them to get their hands dirty and find out what they might be interested in; its a long life.
Terrible advice. If I were 17, genetics and bio tech at the next frontier, with opportunities to be more than another corporate drone. AI is a lot of bureaucracy and nepotistic who ya know already.
Let’s make the xcancel link the actual link, please!
Neural net maths are hardly above scientific high-school level.
Somewhere in rural America is a 17 year old that doesn't even have working plumbing in their house still.<p>I'm sure they'll get right on powering up their computer from the hamster wheel, Paul.<p>My heart goes to all the kids out there that didn't get the fair shake let alone fair access to tech that gets these condescending "learn to code/learn to LLM" bootstrappy talks from rich pricks that don't know what life really can be like for a lot of American kids out there.
[flagged]
You can learn how to build an LLM from scratch while doing all of those things...
Is going to university really that good of advice nowadays?<p>Everybody goes to college nowadays and the average white collar has lots of debt and relatively minor financial benefits over a skilled trade worker.<p>edit: woah, so many people insulted by that. In my bubble and friends, me and another friend are the only people that make very good money compared to non-graduates. Plenty of others opened their shops, went into trades, one learned to tattoo fake eyelashes, one became a (successful) farmer and most make significantly more than the average law/chemist/mathematician/physics/architecture/languages graduates. Sure, the lowest salaries are to be found among the non-graduates too, but I don't see any evidence that graduates make that much more, and that graduating is worth it.<p>Some answers talking about how "formative college is", but my 25 years old friend with her own shop knows more about real life, business and economy than ivy league MBAs.
Depends on where you live (lots of countries have no tuition fees or far lower than the US), what funding you have, how good a university you go it, what you want to do (some careers require a degree), whether you will enjoy it, and whether you are there just for financial benefits or more than that.<p>Its not good advice for everybody, but it is good advice for a lot of people. What if you want to be a doctor? What if you want to work in R & D? Not everyone enjoys working in a shop or a farm. Also, how old is your friend group? If they are mid twenties you are ignoring the greater scope for advancement in a lot of white collar careers.<p>> my 25 years old friend with her own shop knows more about real life, business and economy than ivy league MBAs.<p>Within the narrow limits relevant to her business. How much does she know about macro-economics or financial economics, or scaling up a business? I also suspect you are comparing her to people who went straight on from bachelors to MBA (which is a bad path - study business after having some experience IMO) and lack experience. How will she compare in 10 years time when those people also have real world experience?
If you’re talking purely financial benefit: yes, it is still good advice, as most white collar jobs still require a degree.<p>If you’re talking about a place to mature, around others who are at a similar phase of life, also yes.<p>It is where most people meet their cofounders, for example, even if they don’t found anything until much later.
But it is not true financially, it really depends.<p>Plenty of blue collar workers make more than white collar ones, and have a huge debt free head start in life.<p>A plumber or electrician will make significantly more than the average bank employee or translator or nurse or teacher.<p>And they will also have an easier time starting a business as many trade workers are self employed and make much more than hired ones.
Going to college is worth it for basically everyone other than those that are choosing between a mediocre college / mediocre degree and a blue collar route. AND if college is going to put you into debt.<p>Financially, intellectually, socially, it’s a good idea - college is a formative period of life in American culture. That is more and more true the better the college gets. At the upper tier (Ivy League, etc.) you don’t really have to pay anything if your family isn’t already wealthy, and the connections and degree you’ll make more than pay for themselves.
It depends on the university and your goals, I suppose.<p>Many people are surprised to learn how affordable elite colleges are if you genuinely need financial aid. I had no idea -- was pleasantly surprised when my alma mater took over 80% off of my tuition.
> Is going to university really that good of advice nowadays?<p>Just some anecdata, but every single one of my university professors was quite bad, but they think that since they are the professor, that means they are smart and the expert.
[dead]
So get into debt and spend money is the advice?<p>Then vote for someone who will make the debts go away?
As opposed to what, neglecting the human experience to grind yourself to the bone for those who own capital, and then voting to uphold that capital? Makes no sense.
This also somewhat accurately describes what a VC backed startup does.
> Then vote for someone who will make the debts go away?<p>As opposed to vote for someone who gives even more power and money to the oligarchs? You bet your last dollar that people will choose the former over the latter.
yes?<p>I mean, have you seen the options for people graduating right now? How people are behaving?<p>Or forget the data, look at how the story of the new future technology is being told. The people making it recognize that it has the potential to put swathes of white collar workers out of jobs, and they are openly talking/warning/PR-ing about it.<p>People in <i>tech</i> and SV, the places which have a underlying culture of near delusional optimism, are talking about trying to avoid being part of "the permanent underclass".<p>Gambling is up, and prediction markets are being treated as financial investments. Wall street bets is a thing, and outright speculative investments are the hope people have to get ahead.<p>This is happening in the USA, forget the weaker or smaller economies.<p>When people see the future as one massive zero sum game, with no way to win by building, then they are going to change how they plan their future.
[dead]
If I were a billionaire, I'd learn how to give away a lot more of my wealth.
This is always such a nonsense question-answer thing, asking a person who already succeeded what they would do if they were young.<p>Even worse when they ask themselves.
Venture capitalist suggests everyone to become his future employee, just as he has been (successfully) doing for his whole life.
Can't this guy enjoy being rich in silence? His takes get worse with every passing year.
I genuinely don't think telling young people to do anything tech-related is good career advice. We don't even know if entry level roles will ever come back. The situation couldn't be worse for these roles. PG thinks that some random teenager will build a startup and get rich from it, or some shit. Like get real, man. Completely out of touch, tech bro who hasn't worked a real tech job for the last like 20 years... Remind me how many startups succeed again, Paul? What about the market dynamics for LLMs and how one might "secure compute"?
If I were 17, but have the money I have now he means.
The times are different. When I was 17, we had to buy records; a 17yo today can listen to the whole of the available produced music, plus interviews and all other uncommon and related material, for free from the comfort of "here and now".<p>Possibilities exist now that did not exist before. Those who do not exploit this are fools.
What the fuck does this guy know about? I'm sure if we went back through similar statements he's said over the years he's said the same thing about various technologies that are no longer relevant. The guy is a talentless hack who larps as a blogger and his only "redeeming" quality is having lots of money.<p>Owner of Golf Club Company says I should dedicate my life to golf lmfao.
hackers and painters and kids and ROI and startups and capitalism and the destruction of nature and old guys with money talking like they know better in fascist social networks
The amount of people who missed the point here is absurd. He's advocating for learning about how LLMs work. For the sake of learning. Because no one's going to invent the next thing without at least some understanding of the current thing.
I think collectively we should all stop listening to Mr Graham..<p>He capitalizes on greed and hype but with a soft, sober and thoughtful voice so as to lull you with rationalism and now 20 years of his “disruption” has mostly ruined modern society and a whole generation of techies have been led astray into trying to “change the world” is the world of today (minus the magic technology really any better than 20 years ago?)<p>- Good for him and his Tech Bros, bad for the rest of society
Twitter mirror: <a href="https://nitter.net/paulg/status/2091544343589060625" rel="nofollow">https://nitter.net/paulg/status/2091544343589060625</a>
[dead]
[flagged]
[dead]
[dead]
[dead]
I do not think it is a proper thing to do for 17 y.o., unless they are exceptionally mathematically gifted, as proper understanding of how LLMs are trained requires a good grasp of calculus, understanding modern OS and SDE tools for proper implementation of pipeline etc.<p>I'd rather simply write another mnist implementation and check if I really like all that AI stuff at first place. Even then, before going into mature-on-the-way-to-dying tech (LLMs) I'd rather focus on fundamentals - good ols linear models, regressions, stat etc.
> I do not think it is a proper thing to do for 17 y.o<p>If I'd get a buck every time someone said something like this to me when I was in the 13-18 range, I wouldn't have a ton of money, but it's so very annoying when people tell you this.<p>Regardless if they're "gifted" or not, regardless if you believe in myths like that or not, let children explore what they want to explore, even if you don't understand what it is or why they want to explore that, just let people explore, regardless of age.<p>It was such a terrible experience being a young kid growing up, with so many adults spending hours trying to convince me to stop sitting in front of the computer so much doing whatever; "why are you even trying to learn that stuff, you have to go to school to understand anything of this" and so much other similar trash.<p>Sorry, not your fault and I'm borderline trauma-dumping now, but really sad to see this sort of gatekeeping on HN of all places, age is irrelevant to learning ANYTHING, in my humble opinion at least.<p>Kids, find anything interesting? Jump into it, ignore what adults tell you, and do whatever you feel like, you'll find your place eventually.
I just voiced my opinion. I just think buiding an LLM from the scratch for 17 y.o. is pointless exercise, advising a teenager to do so is borderline irresponsible, and frankly PG is simply virtue signalling here, as LLMs are still trendy, esp. in his circles.<p>There still will be varyy small number of outliers among youngsters who'd be able to extract tremensous value from such an excercise, but for most that'd be _IMO_ waste of of time, with illusion of understanding w/o actually having any.
> I just voiced my opinion.<p>Same! I just happened to disagree with your opinion, and frankly, I'd say trying to gatekeep what people learn is closer to "borderline irresponsible" compared to asking people to build/learn/do X.<p>> youngsters who'd be able to extract tremensous value from such an excercise<p>But they're youngsters, who are about "extracting value"? Life is about fun, not extraction, not value, not avoiding waste of time but literally enjoy what you do, nothing is more important (IMO).<p>Then who knows, doing fun stuff sometimes lead to useful stuff, like in my life. But if you only think about "extracting most value for time spent" or similar "optimization strategies", then you'd never discover this part of life.
> Life is about fun, not extraction, not value, not avoiding waste of time but literally enjoy what you do, nothing is more important (IMO).<p>This is, pardon, demagoguery. There is always "future fun" and "present fun" which a normal person would assign different nonzero weights (<a href="https://en.wikipedia.org/wiki/Discounted_utility" rel="nofollow">https://en.wikipedia.org/wiki/Discounted_utility</a>). Besides, building a LLM _truly_ from the scratch, just using the famous 2017 paper and numpy manuals is not fun at all, esp. for a high schooler.
> Besides, building a LLM _truly_ from the scratch, just using the famous 2017 paper and numpy manuals is not fun at all, esp. for a high schooler.<p>To <i>you</i> it isn't, is my entire point here. But why extrapolate what you think is fun, to others? Sure, I don't find that fun either (although useful), but who am I to say it isn't fun for <i>others</i>?
We can continue this pointless conversation, in the tone "who you are to tell what is fun to others and whst is not". You'd be impervious to any argument stating that dealing with far beyound someone understanding and requiring countless hours of digging into difficult math is not fun even for those who thinks it should be fun, as they presumably, loves everything STEM.
Agree with this - started programming through learning scripting in ROBLOX when I was like 12 (this was back like 17 or 18 years ago) and it developed into a life-long passion for software engineering. I am thankful I had people around me (my parents), who were aware enough to realize I wasn't just playing video games and gave me the time I needed on the computer to learn and experiment with programming...<p>This also meant that by the time I was actually offered to take a programming class in school (junior year of HS), I had already been able to self-teach myself well beyond what that class was covering, thanks to just working on random projects that scratched an itch I had at the time, looking up anything I didn't know or understand, and internalizing those concepts over time.<p>In short though, I definitely agree, young kids and teens (and also, frankly, adults too!) should be encouraged to explore things that they have a passion for, without being told 'you need to go to school for this' or 'you cant understand this at your age'
I think he is basically saying YC has enough startups that are just making API calls.
Why would you tell people that the correct order is to build foundational knowledge before exploring a subject? For some (many?) people, a 'proper' understanding develops _after_ the exploration.
> Why would you tell people that the correct order<p>Because I can?. JK. Because that was my experience, of someone who is 2.5 older than 17?<p>> For some (many?) people, a 'proper' understanding develops _after_ the exploration.<p>I am afraid you have a too confrontational attitude here, but I'll answer anyway: because I do not believe you can simply "explore" such complex topics like building an LLMs. You'd simply be unable to build LLM drom scratch, unless you'd call cargo-cult chaining magic numpy incantations you've taken from Karpathy's tutorials "exploring".<p>If I were in "exploratory" state of mins, I'd rather go from entirely different side - I'd try playing with LoRA-ing existing small LLMs, such as venerable 2 y.o. Mistral Nemo, to get "feeling" for what training is and how hyperameters influence the process.
> I do not think it is a proper thing to do for 17 y.o., unless they are exceptionally mathematically gifted<p>I attempted many projects at a young age that I was absolutely not equipped for. The result of the attempts more often than not left me equipped, every time it left me better off. This is terrible advice.
That'd would be a terrible advice if there weren't a plenty of other things "you are not equipped for", but far less daunting both theoretically and practically. Such as, say, convolutional neural networks, or some older ML tech. Or even something totally unrelated to ML.<p>Transformers are difficult to understand even to people with strong ML background, let alone a teenager.
You are assuming the 17yo in question as an untrained underdeveloped savage. If I were 17 in 2026, I would certainly have exploited all the availabilities from 2010 on - including YouTube, OpenCourseware, the Web simply (Sebastian Raschka etc.) and LLMs.<p>That 17yo would have already built many uncommon bases, and would build further.<p>That is, a 17yo with proper mentality.
>understanding modern OS and SDE tools for proper implementation of pipeline etc.<p>Can you provide an example?
Another great quote by PG. I've been really enjoying his essays recently - truly a great and curious mind.
If I were 17, I wouldn't be using a social media plattform run by racist neo-fascists...
He bases this decision on all of the experience he has amassed, as a 61 year old man in the tech industry. An actual 17 year old, with 17 years of experience, would not think like this, nor should they.
Naturally 17 year olds don't think long-term like this which is why PG's advice is so useful. It gives them a pathway to follow that they likely wouldn't have reasoned otherwise.<p>I told my much younger brother when he was 12 what programming was and it'd be a great career. He looked into it and within months was writing CLI games. Eventually releasing his own unity 3d game on steam as a teen.<p>Eventually he got into CS and did really well because none of it was scary and new. He parlayed that into role at Meta out of university.<p>My point being, 17 year olds have time to learn new skills and guidance can go a long way.
Yes. And they almost certainly have a better understanding of their own situation that him. This is not a dig at Paul Graham, the closer anyone is in age, the better they understand what they have to deal with. I'm roughly in the middle between Paul G and the 17 year old, and even though I'm really quite fascinated with zoomer culture and probably come more in touch with it than most (due to relatives in the age range etc.) I realize I have very little idea what it's like to grow up in the world they grow up in.
Most. But, that isn't the point.<p>Either way, this isn't really advice for 17 year olds. Pg is thinking out loud about the pathways for founders.
What about a 19 year old?<p>>Whoa. I’m 19 and I trained a 100M language model from scratch. Did a v2 now with a new SFT experiment to see if I can get better results on same size.
This is the ultimate builder’s mindset. Tinkering with the hardest technical problems—even for frivolous things like games—always yields the highest return on curiosity. Time to go back to the fundamentals.