49 comments

  • frereubu 1 hour ago
  • nsainsbury 1 hour ago
    It's actually sad how much AI has either not improved life at all or actively made it worse from the perspective of the average person:

    Your car is still the same. Your dishwasher is still the same. The train you take to work is still the same and never comes on time. Fuel is more expensive. The roads are still congested with traffic. Your kids are (probably) doing worse at school. Food costs more. Houses cost more. Rent is higher. Buying a computer or a PS5 is more expensive. Politics is still full of mentally unstable people. The environment is getting worse. You will still die of heart disease or cancer. Food quality is worse. Your kids can't find work. Wealth inequality is accelerating.

    But hey...on the flip side...a lot of people are rapidly building software (that nobody is using), and AI is solving mathematics problems (that nobody - including mathematicians - wanted it to solve)

    Imagine what a "country of geniuses in a datacenter" will be able to do? Apparently...nothing.

    • trescenzi 1 hour ago
      This is the core hubris of everyone behind the LLM craze. I do not want to deny their utility there absolutely is some there. But the inherent assumption behind the whole thing is that with a "country of geniuses" problems will melt away. But that's not the case. The reason problems like housing or famines or war aren't being solved is because structurally there's no incentives to solve them not because there's not enough geniuses out there.
      • redwall_hp 20 minutes ago
        And the LLMs exist to further those problems. People can't afford houses? Well, now they won't be able to own computers either. That's a feature, not a side effect. Once the K-shaped economy progresses enough, you won't even be allowed to use the LLMs. They'll exist only to be a tool for the new feudal class who owns you.

        Unless you destroy their wealth now.

        • peddling-brink 4 minutes ago
          > That's a feature, not a side effect.

          You’re implying that the deep lizard state planned the LLM craze?

          Sama might be a lizard..

        • azan_ 14 minutes ago
          Once capabilities catch up with the demand (~which is in few years - it just takes time to build!) average person will have much better access to higher quality and cheaper hardware. It's literally how it has always worked, I don't understand why would it be different this time. There's zero substance in your comment, just repeating tired tropes from twitter.
          • vrganj 1 minute ago
            A market requires scarcity, because scarcity is what neccesitates trade.

            You can't invoke market mechanism to defend a post-scarcity scenario.

            Why would the ai owners make things cheaper or even available at all once their abundance machines give them everything they want?

            What would they get out of if it?

            You can't argue "this will change everything" and "this is how it's always worked" simultaneously and that is the core oversight behind AI boosters.

      • aleph_minus_one 57 minutes ago
        > The reason problems like housing or famines or war aren't being solved is because structurally there's no incentives to solve them not because there's not enough geniuses out there.

        Rather: these problems are not solved because there are not enough geniuses who are capable of convincing the masses to send the responsible politicians to hell.

        • techblueberry 31 minutes ago
          In America anyways, we send the “responsible politicians” to hell every 2-6 years. People vote for this.
          • plasticchris 10 minutes ago
            and yet... it feels like we are actually a one party state most of the time.
            • Avicebron 1 minute ago
              The party of corporate interest?
        • blactuary 44 minutes ago
          A great deal of the supposed smartest people have fully sided with fascism, it's not about smarts
          • majormajor 32 minutes ago
            This is broadly the same people who bemoaned the idea that the "smartest people in a generation" were all just working on adtech without any reflection about maybe that's not actually an exclusive set of the smartest people... or that a lot of the people in that set were the same one who might, twenty years ago, otherwise be chasing money as middlemen in finance vs doing productivity-expanding things anyway...

            You can be smart and chase the money, but it's crazy hubris to assume that only your peer set of money-chasers are smart.

            It's a small leap from there to "since we're the only smart people, we should get to have the most influence."

          • dd8601fn 36 minutes ago
            I’ve heard people say variations of, “Clever people can easily convince themselves of just about anything, if they’re inclined to. And then it’s nearly impossible to talk them out of it.”

            Kinda seems true. And the worst people in history often don’t seem super dumb, despite their various other defects.

          • crooked-v 36 minutes ago
            "Supposed" is doing a lot of work there. Compare, say, Yudkowsky et al. to the people who actually changed the world with mRNA vaccines.
            • jeremyjh 7 minutes ago
              Yudkowsky is hardly an example of someone who has sided with the fascists.
        • azan_ 13 minutes ago
          > Rather: these problems are not solved because there are not enough geniuses who are capable of convincing the masses to send the responsible politicians to hell.

          People support policies that make housing much less affordable (e.g. rent controls) and oppose policies that would make housing more affordable (e.g. zoning reforms). Saying that it's evil politicians that cause this crisis is dishonest - it's what people want.

        • card_zero 37 minutes ago
          For that you might need the kind of geniuses who write stirring books about principles and values. Or whatever the modern equivalent of those books is. Maybe nobody would pay attention to that kind of advocacy in the present day, not sure why. But it's not tech geniuses, anyway. Much though tech has power as a catalyst.

          I suppose there was Project 2025, and there was that book about the inevitability of Russian hegemony (or empire or something) that Putin liked. Those got attention with influential people. Not sure what to conclude from that.

      • skew-aberration 24 minutes ago
        Yes, many of the problems are 'features not bugs'. There's no reason to think the AIs won't dream up more of these'features' too.
      • Terr_ 53 minutes ago
        To add insult to injury, we're being told: "Just trust us bro, however we make things worse now will totes be redeemed as the prophesied god-mind magically fixes it later."

        Strangely, they never discuss the possibility that Multivac might seize all their wealth and cast them down into average-ness.

    • TheAceOfHearts 55 minutes ago
      I'm immensely sad every day seeing all these posts talking about billions and trillions of dollars and where I live (part of the US btw) we don't even have fully reliable electricity and water access. In fact our electricity bills have increased and there is nothing to show for it.

      I don't understand why these big labs aren't doing more to improve the electricity situation. Dump a billion dollars worth of tokens into fusion analysis, research, and simulations and try to force a breakthrough. You could buy a lot of good will with unlimited cheap energy that liberates us from fossil fuel dependencies.

      • ronsor 51 minutes ago
        OpenAI and Sam Altman are literally backing a nuclear fusion power startup.
        • blactuary 42 minutes ago
          "Startup"

          Wait until you have clean energy sources and shut it down

      • ungovernableCat 42 minutes ago
        Mathematics and programming is verifiable and fully controllable in a digital environment.

        Fusion research and cancer research requires expensive experimentation in the real physical world. You can’t just throw a billion dollar worth of tokens to handle the problem.

        Look at all the roadblocks for data centre buildouts which the industry certainly has a financial incentive to complete as fast as possible.

    • atleastoptimal 22 minutes ago
      >AI is solving mathematics problems (that nobody - including mathematicians - wanted it to solve)

      Do people need permission from mathematicians to solve math problems? Is one of the Millennium problems not worth applying a new system to work on it?

      I understand the despair people are feeling at AI automating their work, but we have to be serious about the fact that the frontier labs are just responding to natural incentives. Mathematics, coding, all are just applications of logic, and it makes sense that a system trained on inferring patterns from data would work well in that domain.

      • jeremyjh 5 minutes ago
        You are misunderstanding what mathematics research is. Its objective is the development of human understanding of mathematics. This requires open problems to research. AI are “solving” the open problems without creating any human understanding.
    • kozikow 1 hour ago
      At least for me, personal agents have been quite useful. If you integrate them well - shopping, filling random government forms, keeping track of my personal life, optimizing training plan and nutrition, planning vacations, etc. I can do more home repairs or renovations myself with help of chat

      So basically you get personal digital assistant - very useful - but not changing my life.

      What would I need for real transformation though? Food, house cleaning, commute - "every day chores". Even if we had ASI tomorrow, what could it do to change my life, besides taking away my job? Design good house robot to help me with food preparation I guess... Maybe finally self driving cars launched in my area? Or it will take control, "optimize humanity survival chance" and realize we don't need as many people on this planet.

      • ptx 59 minutes ago
        Almost every time I use an AI chatbot it eventually turns out it's lying to me. I would not trust them with any of the things you mention.
        • Aurornis 38 minutes ago
          This is the kind of gap in shared understanding that is widening.

          When people use words like “chatbot” to talk about their AI use and can’t really name basic things like the model they were using, it reveals a wide chasm between the people who are deep into understanding how to use these things and those who poke at easy options every few months and then look for ways to confirm their priors instead of trying to understand how other people are getting value.

          • majormajor 27 minutes ago
            The incentives seem to be lined up to push users to the frontier of "how much can the harness around the model do and claim without most people pushing back on the claims."

            The bigger the pile of stuff you have the tool create unsupervised, the more incomplete/inaccurate the overall picture of it is gonna be. That doesn't matter for most hobbyist/helper-tool/game-port/etc type of tasks.

            I get a ton of value out of the tools and yet part of that value has come from being hyper-aware of the shape of the limitations, which changes much less and more slowly than the specific of the limitations.

            "How'd you notice that one!?" - well, it helps to be the only person on the team who took the time to interrogate the e2e interaction of all the code...

            The median piece of OSS software was already pretty rough once you had to depend on it in production at scale to make money. The median piece of the new AI-generated wave of software, also pretty rough.

            The bet that OpenAI/Anthropic is making is that they can automate the interactions generally/globally enough to work around this and eliminate the need for humans in the loop. But it's a tricky problem because it's precisely where the permutations and special cases get nasty and much harder to deal with than more broadly-applicable cases of codegen/test-gen/mash-it-until-the-tests-pass-and-the-code-review-angent-is-happy.

          • Madmallard 34 minutes ago
            using skills with opus 5.5 and agents for each individual task that have shared knowledge of all of your records does not lead to a significant increase in trustworthiness
        • satvikpendem 48 minutes ago
          I have ChatGPT Pro and it sources every single thing. For my Asia trip right now it's been helping me every single day to figure out optimized routes on what to do and where to go (if I feel like it, not to say I don't wander around too).
          • blactuary 43 minutes ago
            That sounds terrible. I like to live like a human being
        • heaney-555 47 minutes ago
          When is the last time you used an AI chatbot, and do you only use the free versions?
        • kolinko 43 minutes ago
          Which model? Super weird
          • qwerpy 14 minutes ago
            Every time I use Google search AI for anything more complicated than a search query I feel that way. A paid model from pretty much anyone is much better.
        • kozikow 52 minutes ago
          > Almost every time I use an AI chatbot it eventually turns out it's lying to me. I would not trust them with any of the things you mention.

          If you talk to free chatgpt, sure.

          But if you launch Opus 5.5 in claude code - for some reason Opus is the smartest for me, even for "general" type of questions, when launched from Claude Code. It has harness/prompts optimized for correctness, while more "customer-grade" products, even when used with top model tend to optimize for latency. And I tend to give "context" - e.g. I have "personal wiki" + skills for lots of those things.

          I don't recall i last few weeks where Opus 5.5 in CC made an obvious lie, made an unreasonable assumption, or made a mistake that smart human couldn't make. E.g. if it "failed" it failed in a way that non-genius human could fail in similar way.

          • Madmallard 28 minutes ago
            opus 5.5 uses agents and skills in chat sessions as well and maintains internal context that it tries to distill from what you provide

            I tried to have it write a synopsis website for the different types of dance patterns in Dance Dance Revolution, with guides on their execution, explanations, and provided song sources. I provided all of the source articles and search websites needed for it to be able to do that, and gave it a sample set I wanted it to utilize and build off of.

            It has bits and pieces of information wrong in 90% of the sections it created. Poor song choices, missing song choices, incorrect explanations, incorrect step-patterns, the list goes on.

            If you're not an expert in a given field for which you're using it, it's going to look like a genius. If you are, it's untrustworthy and borderline useless.

      • grtt1 40 minutes ago
        [dead]
    • robswc 1 hour ago
      This is sort of my take. You could look around in 2016 and 2026 and honestly not much has changed... at least at first glance. There is a missing piece of the puzzle.

      I'm noticing this even in industry. Yea, we can build and iterate on ideas 100x faster but has this turned into anything meaningful? Not really... Hell, look at the big companies spending billions on inference... nothing shipped and things are still buggy as ever. Where is everything?

      • redwall_hp 12 minutes ago
        Sure, it did some meaningful things:

        * Higher load on senior engineers who are about to burn out hard from fighting an endless battle against unreviewed slop people keep trying to yeet into production

        * Less junior hiring and suppressed wages

        * Culture of uncertainty and fear

        * Computers are really expensive now

        * Some people now live with diesel generators and gas turbines making their neighborhood hazardous to their health

      • Rapzid 1 hour ago
        The job market had changed that's for sure.
        • zer00eyz 47 minutes ago
          We went from spending a lot of money on people to enshitify products and then charge more for them to spending a lot of money on agents to enshitify products to charge more for them.
      • aorloff 18 minutes ago
        That's what I'm saying.

        Not only from the big companies. Where is the explosion of disruption ?

      • prometheus1992 59 minutes ago
        >You could look around in 2016 and 2026 and honestly not much has changed

        I find 2026 much more repulsive compared to 2016.

    • satvikpendem 50 minutes ago
      This is a Western-centric viewpoint. I'm in Asia visiting a few different countries right now, including the big ones, and much has changed with AI and also in general over the past 10 years with cars, trains etc.

      > AI is solving mathematics problems (that nobody - including mathematicians - wanted it to solve)

      Strange viewpoint. In mathematics any novel result is useful as it expands the scope of other results. Who are these mathematicians who didn't want AI to solve the problems it did? I see trepidation about the future of the field but that is not equivalent to not wanting them to be solved.

      • SyneRyder 16 minutes ago
        Terence Tao and the apparently newly formed Association For Human Mathematics. The post is online-famous now for the line "Mathematicians did not ask for this work to be done."

        https://terrytao.wordpress.com/2026/10/07/ahm-statement-on-o...

        • plasticchris 5 minutes ago
          When I was a young man reading gauss, and thinking before I turned the page, I see a result here, something meaningful! and then I turn the page, and of course it is there. There is no shame in this.

          A long time ago, another student of mathematics described it to me as climbing a mountain. No single step is hard, but you must make many steps to get to the point that no one else has made a step.

      • AnonymousPlanet 29 minutes ago
        I'm curious, care to name a few things that have changed in Asia with AI (and wouldn't have without)?
      • m0llusk 16 minutes ago
        That is not how math works. There are reasons why the Riemann integral is pervasive and the Labague integral is rarely used unless needed. Just because you can do something does not mean that you should.
    • c0rruptbytes 59 minutes ago
      LLMs could’ve been developed a lot slower without capsizing the entire supply chain and been seen a lot more positively, China is building a lot of good AI with a minuscule of the compute power and it’s fine

      there was never a need to rush other than for capitals self interest to control

      if frontier labs wanna pace the frontier, they should be honest and return the compute back to the people

    • mindwok 1 hour ago
      Whatever benefits AI has to offer on a societal level are going to take a long long time to show up, give it some more time. We had smartphones in 2007, and e.g. By way of example, Uber didn't launch ridesharing until 2013, and then that still took years to seep into the public awareness.

      By comparison, we're in the early years of AI. We're all still figuring out what this thing even means.

      • fasterik 21 minutes ago
        I agree with you, but those examples are going to fall on deaf ears. The pessimistic mindset will say that smartphones were a terrible invention that's ripping society apart, Uber is an unethical company exploiting workers, etc.
        • dofm 20 minutes ago
          When did Uber stop being unethical? Did I miss the memo?

          Because they didn't just break laws, they built an entire system to mislead law enforcement.

          (How do tech people not know about Greyball??)

    • raz32dust 1 hour ago
      This view is too cynical. The frontier advancements are certainly going to trickle down. Take cars for example. AI driven cars are already here. AI-designed battery systems, fuel systems, materials and 3d printing capabilities will come to cars. It will likely lead to cheaper, more efficient and more effective vehicles. If will just take time for the impact to percolate.
      • mym1990 9 minutes ago
        If you look at the average price of a new car, adjusted for inflation, it has pretty much always gone up aside from recessions. There have been plenty of advancements in car technology, they do not make them cheaper. It actually gives reason to charge more for them due to perceived increase in value.
        • raz32dust 0 minutes ago
          Yes but you're getting a lot more for the same price.
      • eloisius 47 minutes ago
        My prediction is that while those “advancements” may come to cars, the only practical differences will be lower quality products, cost savings for the car companies, probably more expensive for consumers, and somehow it will also be collecting more data to sell to advertisers.
      • blactuary 41 minutes ago
        Just a few more weeks, a few more hundreds of billions, it's right around the corner....
      • m0llusk 11 minutes ago
        AI car navigation is here, but driving is a social act. Have you ever tried to keep an automated taxi out of a private driveway or the scene of an emergency? There is a lot of work to do and limited evidence of LLMs raising the bar.
      • xingped 50 minutes ago
        Ah yes, the famous "trickle down" theory, which you may know from its smash hit "trickle down economics"! Surely this time it'll be different, right?
        • elmer2 37 minutes ago
          How can you say tickle down economics doesn't work? The US has made it easier for businesses to grow and the result is the highest standard of living ib tje world.
          • mym1990 7 minutes ago
            Hate to break it to you, but the US doesn’t even break the top 10 in a lot of reports.
          • ks2048 18 minutes ago
            I think it's hard to quantify "highest standard of living", but for the attempts at it (e.g. HDI, etc), the US is not number one.
          • ba1afd89f34cb23 29 minutes ago
            Trickle-down economics specifically refers to economic policy that favors the wealthy and corporations with the argument that this will lead to wealth trickling down to the economically less fortunate. Examples include the Reagan and Trump tax cuts.

            This has been studied somewhat extensively (https://academic.oup.com/ser/article/20/2/539/6500315?login=... is one example), and it’s been shown over and over that this does in fact not work and leads to larger wealth disparity, across many countries where it has been attempted.

            It’s possible you’re working off some other definition of the term based on your comment but that’s the generally accepted usage.

      • Supermancho 1 hour ago
        Worse, the view is largely phrased as gaslighting AI. Blaming the cause of particular grievances that coincidentally have been sourced at the same time as the rise of AI, is purposeful.

        > [Your train] never comes on time. Fuel is more expensive. The roads are still congested with traffic. Your kids are (probably) doing worse at school. Food costs more. Houses cost more. Rent is higher. Buying a computer or a PS5 is more expensive. Politics is still full of mentally unstable people. The environment is getting worse. You will still die of heart disease or cancer. Food quality is worse. Your kids can't find work. Wealth inequality is accelerating.

        None of this has to do with AI.

        > But hey...on the flip side...a lot of people are rapidly building software (that nobody is using), and AI is solving mathematics problems (that nobody - including mathematicians - wanted it to solve)

        Plain wrong on both counts.

        > Imagine what a "country of geniuses in a datacenter" will be able to do Apparently...nothing.

        Nonsequitor, probably based on a personal beef. Datacenters, at best, are understood to be for hosting more AI - not for people to somehow live in as an arcology of discovery.

    • denkmoon 1 hour ago
      Using AI for everything digital, I end up going and doing chores while waiting for the LLM to do its thing. What an insane world, where knowledge work (dreary and uninspired as it is) is automated while I am consigned to scrubbing dishes and waiting for the machine.
      • azan_ 6 minutes ago
        Buy a dishwasher, you won't regret it!
    • mattnewton 1 hour ago
      I think this is related to the structures we’ve built in society. It’s lower friction to write software than just about any of the problems you listed.

      It reminds me of the Ezra Klein quote that “OpenAl would need permits to cover its parking lot in solar panels, but it can accelerate into recursive self-improvement, as best I can tell, whenever it so chooses."

      • numbasys 57 minutes ago
        Soon: The super-intelligent AI is telling us to cover parking lots and other flat roofs in solar panels.
    • elmer2 43 minutes ago
      AI has only been around a few years in its current form.

      It might take a decade to take what we learned from AI and turn it into something that will benefit the masses.

      Most of it can't be broken down into something as simplistic as a dishwasher.

      Most people won't notice cancers being found faster due to AI being better at pattern recognition than a doctor, but it's still a great benefit.

    • tim333 16 minutes ago
      I think one of the hopes is they'll get smarter than us and able to solve real world problems better - disease, war, poverty etc. But they are not really there yet. Give it a little while.
    • chiengineer3 23 minutes ago
      I've been thinking about this for 2 years. At some point it will be easier for them to create self sustaining cities - free houses cars public transportation clean food clean water -

      Africa is going to look like wakanda in less than 5 years if WW3 gets canceled

    • api 6 minutes ago
      A good number of the things on that list are things tech can't fix because they're not technical problems. They're political problems.

      We could fix them, but either the political class doesn't care or actively opposes it.

    • emdash 1 hour ago
      That pretty much sums up my experience. Everything is more expensive and worse at the same time. And it feels like we're being screwed over from every direction and in new ways every day. People keep talking about this tech is supposed to do so much but everything just gets worse.
      • fasterik 15 minutes ago
        It's not true that everything is more expensive. Nominal prices are higher, but real income has risen faster than inflation. Purchasing power of the average person in 2026 is higher than in 2016.
      • SamPatt 39 minutes ago
        What time scales are you using to make this judgement?

        Everything isn't getting worse constantly. The data is quite clearly the opposite. How old are you?

    • fasterik 41 minutes ago
      Is it reasonable to expect AI improvements to filter into all of those aspects of life that quickly? It's only in the past 3-6 months that we've started to see major breakthroughs in mathematics. It might take several years for that level of research capability to get adopted in areas like engineering and medicine, at least in a way that impacts the average person's quality of life. But at this point it seems highly likely we'll get there eventually.

      Your pessimistic view doesn't hold up against the data. For a lot of the problems you mentioned, we're in a better place today than almost any other time in history. Of course, things can always be improved and we should never grow complacent or accept injustice. But by most objective metrics, there's never been a better time to be alive.

    • Ericson2314 13 minutes ago
      I mean there's been huge strides in electric vehicles at low prices that Americans haven't experienced.

      We've just locked out a large part of modernity because the shitty American auto industry is in swing states and also gender.

      Same could be said about the slow rollout of Waymo.

      Lots of America is just very anti technological change, and this means any technology no matter how "magical" hits a wall.

    • hellohello2 58 minutes ago
      Publicly-accessible AI started working 1 year ago, 2 years tops. What did you expect? It to solve war driving prices up?
    • ckocagil 13 minutes ago
      Why fix your appliances and infrastructure and give you healthcare when the plan is to obsolete the whole lot of you?
    • oblio 1 hour ago
      https://nitter.xitter.cc/augeeidos/status/210940046382122618...

      This reply summarizes the problem from a different angle.

      • aleph_minus_one 33 minutes ago
        This is an interesting perspective.

        Nevertheless, my counterpoint to

        > If they actually had access to systems that could do "years of top professional work" across software, cyber, science, and math, they would not mostly be leaving to start more chatbot companies, agent frameworks, eval tools, and "AI for workflow" startups.

        > They would be starting normal companies to disrupt the global economies of the world.

        > Insurance companies. CAD companies. Compliance companies. Security auditing firms. Logistics software. Tax automation. ERP replacements. Drug discovery tooling. Industrial design. Legal ops. Accounting systems. Boring vertical SaaS with insane output per employee.

        is:

        Many people who work in such industries already know an insane amount of inefficiencies. Unluckily, they are not in positions and have no political power (with respect to organizational policy) in the company that they are capable of changing the situation.

        Additionally, many of these sectors are heavily regulated, so it is very hard to make a difference in this sector as a startup founder, and even if there is an entrpreneur who founds a company to change the status quo, an insane amount of work in the company will be handling the big amount of red tape.

    • ajross 30 minutes ago
      > But hey...on the flip side...a lot of people are rapidly building software (that nobody is using), and AI is solving mathematics problems (that nobody - including mathematicians - wanted it to solve)

      That's not fair. To first approximation, most of the progress in human condition of the last, like, forever is implementing solutions that people didn't see value in at the time of creation. Woz built a toy that could barely run software at all. The human genome project was panned as a boondoggle. Penicillin was a mistake.

      The first fire pit probably looked like a dangerous stunt to the tribe faced with it, even.

      Does that mean everything will come up roses? No. But LLMs only started talking to us for real like three years ago. Let's not just that fast.

    • heaney-555 48 minutes ago
      >Your car is still the same.

      Not for Tesla owners in select countries, where their car now drives them around 99% of the time, or the people of Austin Texas, that can now hail a fully driverless Cybercab for cheaper than an Uber!

      >Food quality is worse.

      Food quality is higher than ever if you want high-quality food!

      Unfortunately, the people crave slop.

    • TacticalCoder 1 hour ago
      I'll add some...

      Videos are shittier. Most songs are even shittier than the autotuned pieces of shit already were. Ads are worse (even though nobody thought that was even possible). Customer support is worse.

      > But hey...on the flip side...a lot of people are rapidly building software (that nobody is using)

      That nobody is using and that are worse than what we used to have. And because nobody's using them, the countless issues plaguing them aren't even reported by users anymore (there are no users anymore). Full of bugs, usability issues and gigantic security holes (but which doesn't really matter because nobody's using them).

      It's not even clear that it's making long-standing software (like say Linux or Emacs or Blender or QEMU) better at all.

      Now I'd say though: most of the issues you mentioned have nothing to do with LLMs / AI. But it's very clear that AI ain't solving much of the world's problems and that the promises we got years ago didn't materialize.

      • Roark66 50 minutes ago
        All you said is true, but on the flip side it is giving an individual user huge power over software. To create their own, to reverse engineer proprietary ones. Decentralisation of such power can only be a good thing.

        Even if a certain small group of people are also using the hype surrounding this tech to attempt to destroy the home PC market and reverse the personal computer revolution. They will not succeed. There is more of us than them. Even with "thousands of frontier LLMs". A human+llm is going to be more powerful for a long time rather than a bunch of LLMs.

        • blharr 17 minutes ago
          The power isnt really decentralized when you see that the internal AI models are superior than anything you can self host, and the right to use those can be stripped away at any time.
    • ModernMech 36 minutes ago
      Have you seen the shareholder value though? That’s not nothing it’s the only thing that matters!
    • dreamcompiler 1 hour ago
      Correction: My car spies on me and its cruise control no longer works because it slams on the brakes in the middle of the highway whenever it sees a mirage. My dishwasher broke because it got an unasked-for software upgrade.

      Things worked better before we started putting bad software in everything. As a software engineer, I know what good software looks like. The fact that almost all software is now bad -- including the software that drives AI and the software that AI generates -- makes me want to flip the table. We know how to write better software and the general public should hang us if we continue to refuse to deliver it.

      • Terr_ 42 minutes ago
        Probably preaching the choir here, but the software is just a symptom.

        The problem are the power-relationships and laws which allow someone to abuse you and your stuff and then put you in jail for trying to remove the software.

      • drnick1 41 minutes ago
        > Correction: My car spies on me and its cruise control no longer works because it slams on the brakes in the middle of the highway whenever it sees a mirage. My dishwasher broke because it got an unasked-for software upgrade.

        Never, ever connect appliances to the Internet. I thought this was common knowledge on HN. Also pull the telematics (cell modem) fuse to prevent your car from snooping on you and selling your location and driving habits to data brokers and insurance companies.

    • scotty79 1 hour ago
      You are talking like we had this capabilities for a decades but it's been just a few months of truly capable AI. Virtually all of IT switched to agents that quick. Non IT people just started to make software for themselves. Vision and translation capabilities of LLMs help millions. People talk with AI's now verbally to get advice or to just hear a friendly voice. The math thing is just few weeks old. We are just getting started. How fast did the telegraph meaningfuly changed lives for the better for bulk of people? Radio? TV? Computers? Internet? Cellphones?

      Has any technology made an impact on physical world faster?

      • LaurensBER 1 hour ago
        Exactly, we're all software engineers here and even for us it took weeks/months to get to a position where we can really leverage LLMs (think platform, security, observability, etc).

        Imagine being in charge of a physical factory, sure you might have a 200 USD Claude subscription but you also have a floor full of physical machines that have, if you're lucky, an undocumented debug interface and a paper manual. Not to mention safety inspections, insurance, etc.

        You can't vibe code a factory floor. Not yet, perhaps in 5 - 10 years but there are some very hard, non technical, problems that need to be fixed before we'll see meaningful improvement there.

      • oblio 1 hour ago
        > to just hear a friendly voice

        This is far from being proven as a net gain and there are massive lawsuits incoming related to AI recommending suicide to kids and others.

        • LaurensBER 1 hour ago
          In all fairness, depression and suicide have existed before LLMs.

          Safety systems are not optimal but improving everyday and although I doubt we'll ever get a perfect safety system, it seems that we're not very far from a "good" safety system. The days of Google recommending glue on pizza and generating diverse Nazis seem behind us.

          Not to mention all the untold good that LLMs have done. Personally I've seen that LLMs are incredibly good at pointing conspiracy nuts and extremists into the right (moderate) direction.

          In a polarised world, we definitely need that.

    • sans_souse 1 hour ago
      I would just add that consumer products in general aren't the same - practically every business model has incorporated obsoletionism to the point it's practically cancelled out the curve of progress and evolution of tech. We could have nice things.
    • Aurornis 40 minutes ago
      It’s amazing how much the goal posts move in every AI conversation. A couple of years ago the average HN thread about AI would be filled to the brim with skepticism about AI ever being able to accomplish anything other than flappy bird clones.

      Now it’s comments like this lamenting that the amazing LLMs we’ve had since early 2026 haven’t entirely revolutionized the whole physical world and even global geopolitics yet.

      > But hey...on the flip side...a lot of people are rapidly building software (that nobody is using),

      I guarantee a lot of software you use today has substantial AI development in recent releases.

      Even the Linux kernel is getting a heavy input from AI and Linus is advocating for it. Unless you’ve frozen your entire software stack in 2024, you are using software with AI generated code.

      > and AI is solving mathematics problems (that nobody - including mathematicians - wanted it to solve)

      Saying nobody wants those problems solved is flat dishonest. The complaints from some mathematicians are that they and their peers didn’t get to put their names on the papers, not complaints that the problems were solved.

    • sick_of_slop 51 minutes ago
      [dead]
    • WillowWithAWand 1 hour ago
      [flagged]
      • chanakya 56 minutes ago
        "It is not from the benevolence of the butcher, the brewer, or the baker that we expect our dinner, but from their regard for their own interest"

        When companies make profit, it doesn't mean we get screwed, nor even that they have a high profit margin. The grocery chain stores are a prime example. They typically make 2-3% net margins on revenue, and provide fresh, clean, reasonably well-tested groceries close to our home for that margin. Imagine trying to get equivalent quality fresh groceries ourselves from the producers, and I'm guessing it would cost at least three times as much.

      • tim333 35 minutes ago
        Worst system, apart from the other ones.
  • rico735 56 minutes ago
    Don't think the 0.2% really are in agreement that "large, complex projects that used to take them weeks/months can now be completed by agents with a prompt".

    The divide seems to be between people who think you may need to understand the codebase for maintenance, potential future token costs, optimisation, code erosion & being wary of the assumptions/mistakes that they have seen LLMs make, and people who don't review code and are confident that the AI has made no mistakes and will be available for next-to-nothing forever.

    It's certainly allowing the latter to move fast... but will their clients be happy if/when they break things. We'll find out soon, I'm sure.

    • SamPatt 32 minutes ago
      It's not that we're assuming the AI has made no mistakes. It's that we're confident in the general capability of the models such that we're willing to accept it'll make some mistakes, but working together with it, we'll still achieve our end goals.

      I presume you also know your own code has mistakes, too? You wouldn't refuse to even look at the code of a colleague who has ever made a coding mistake, would you?

      • rico735 0 minutes ago
        That's not really my point... my point is that it's an unsettled debate between speed and caution, and devs are not unanimous on it, unlike the claim of the original post that seemed to think everyone was reporting the same gains.

        The divide with developers at the moment seems to be about people who believe in review & are being cautious (who are not seeing major time gains according to all reports I have seen), and those who trust it to be able to fix things forever and will no longer have to worry about understanding the code.

        Yes, my code will have mistakes, and I'll use LLMs for some tasks. I also review and have a mental model, which means I can easily fix them, no matter what happens in future with AI. Reviewing means I'm seeing limitation.

        A no-review approach will be far faster than this, obviously, but if scbench.ai is still true on the frontier models, it may come back to bite people when they have no idea how it works and the context becomes too big to maintain.

    • Eezee 32 minutes ago
      I can do a project that took me weeks/months much faster now as long as I don't have to involve other people in the decisions.
      • mrbombastic 14 minutes ago
        I feel like this was always true to an extent, now product and design are just more tolerant of it because leadership is cracking a whip over their heads because things are supposed to be much faster “because ai” so things like pixel perfect hand offs, adhering to existing design system, making sure you are aligned with the brand, copy is reviewed with marketing go out the window. Now everything is shipping in hackathon mode with an eng driving decisions they really should not be.
      • ttul 18 minutes ago
        I should frame this for the wall behind me when I run meetings with my team. It’s the core problem with AI in coding.
  • NichoPaolucci 4 minutes ago
    What exactly do we want the general public to "do" with these models? If the have no use for agentic LLMs, they obviously aren't going to be interested in the capability of those agents.

    I see a lot of benefit to spreading information around the tech, but it's still in its infancy. Most people weren't on the internet when that was getting up to speed.

    Meemaw and pops have their systems that they work within, and they just got familiar with the damn smartphone. Though it might help them, they probably do not want AI to organize their old photos or "clean up" their computer. Saving 5 minutes just isn't important to them.

    As for professionals, yeah - companies are going to adopt these tools over the coming years. It's the "norm" in my circle, but lots of jobs just... don't use technology like that - there has ALWAYS been a major gap between people who are deep in tech and people who are not.

  • purplepatrick 5 minutes ago
    Yes, and this is completely normal. Not sure I get the point of his post. Most likely, stating the obvious is what you gotto do to keep your followers "engaged".

    The average person does not have the time to occupy themselves with a new thing to see what it can maybe do for them if they combine it with hundreds of others. Until non-chat-bot AI comes in discrete usable packages, mass adoption will be limited, because it's all just a big bunch of lego pieces whose utility is hard to grasp, because it isn't self-evident.

    This isn't new, though. It's how "technology" usually works. For example, human(oid)s have been able to stick sticks into things for ~500k years, but the fork wasn't "invented" until ~2-4k years ago.

    AI is a background technology, a building block. Something that is best appreciated when you don't know it's there -- when it just does a job (well). The end-in-itself POV of the tech industry is self-indulgent and irrelevant for the average person.

  • turzmo 5 minutes ago
    It is always a mistake to assume that because computers can do X, that they will be able do everything that a human that can do X is expected to be able to do.

    Solving frontier math problems is mind boggling, but perhaps it says more about frontier math problems than it does about LLMs. Same with chess which was solved decades ago.

    My experience is that LLMs can be useful but are underwhelming in terms of intelligence, namely the ability to quickly learn from experience and grow with you while you accomplish a task.

    Thinking of them as beings is something we should move away from.

  • kooi 1 hour ago
    The 1%ers are in the vertical, but the question is vertical to where?

    It needs to a potential field with practical, economical, "real life" attraction well. I.e, robotics, real economic efficiency gains, manufacturing novelties.

    The worry is that the 1% is attracted towards a non-practical money hole. I.e: Token burn for the lols, sophisticated software systems that dont provide actual value outside of giving NVIDIA cash.

    • freecodeio 1 hour ago
      he's just trying to say we're in a cool kids club and you ain't in it, without specifying what the cool is about
  • ilovecake1984 1 hour ago
    I’ll say this until I am blue one the face. Nerds (software dev, maths etc) see how good LLMs are at things they care about and assume they will be broadly applicable in future.

    There’s no reason to think this.

    • howunfortunate 1 hour ago
      > There’s no reason to think this.

      There are lots of good reasons to think this.

      The first is that LLMs used to be bad at each of these nerd things, then toppled them like dominos. There's a pattern over time.

      Another reason to think this is g. Across every known measurement, human (and animal) intelligence is convergent. Being better at one thing correlates with being better at another thing. Although the reason is not perfectly clear, the pattern is well established, and seems to apply to LLMs too - GPT-6 is smarter than GPT-3 at everything, not just math. There's no reason to expect different for GPT-9.

      • throwawayy6767 1 hour ago
        That's just cocktail party level of thinking. LLMs are getting better at math, code and logic (and marginally better at science and general knowledge) because these are domains that can be objectively verified and thus there is a potentially infinite supply of 'facts' to generate and train on. These are very powerful but ultimately very abstract domains. For everything else the messy real world and its physical bottlenecks gets in the way and there's little reason to expect progress to accelerate. It still takes months to get mice to reproduce and run experiments on, no matter how knowledgeable about biology the models have become.

        If anything, in some domains frontier models have become worse - claudisms and chatgpt idioms are making them notoriously bad at prose without a considerable amount of prompting and tweaking.

        • atleastoptimal 18 minutes ago
          Is good writing verifiable? I don't think it is, but LLM's have been hillclimbing writing quality. That being said this is with the help of RLHF.

          However there are many other domains which have verifiable rewards in the process of learning them, despite their overall impact not being verifiable. For example, the life sciences, an LLM could be given access to data about an organism, and then make predictions about how a drug or gene therapy will affect that organism. In economics, LLM's could create models of behavior, evaluate predictions over time and see how well those predictions match reality.

          >claudisms and chatgpt idioms are making them notoriously bad at prose without a considerable amount of prompting and tweaking.

          I think this is because a lot of people genuinely like the claudisms, even though a small minority of technical people don't.

      • flecomet 1 hour ago
        I don't agree that they get uniformly better at everything, especially subjective skills.

        For example, they have become more and more unintelligible when you ask for explanations or descriptive text. They assume you see the same context as them and shortcut explanations.

        I've had to craft a skill to get them to produce remotely understandable explanations of even mildly complex/non-mainstream subjects.

        This may be Curse of Knowledge https://en.wikipedia.org/wiki/Curse_of_knowledge on their part, and it also impacts human experts but still. Becoming better at one thing does not mean you become better at everything else, although I do agree with you that the better they become the more things there are they become good at, but their ability is still quite jagged and maybe increasingly so.

        • howunfortunate 6 minutes ago
          The delineation between knowledge and intelligence is useful here.

          Knowledge can occasionally get in the way of some tasks, such as a master illustrator trying to draw like a child. Or as you've mentioned, an expert trying to explain to a beginner.

          Even so, I disagree with you about model progress on communication - maybe there's a little jaggedness between minor model versions, but Opus 5.5 is a much much better communicator than, say, Claude 3 Opus.

      • oblio 1 hour ago
        > Across every known measurement, human (and animal) intelligence is convergent. Being better at one thing correlates with being better at another thing.

        LOL. Human (and animal!) intelligence is notoriously jagged. We're all basically idiots except for very narrow areas where we focus.

        • howunfortunate 1 hour ago
          This is simply untrue. I tried to think of an animal example so I could give a concession, and there just...aren't any.

          Smarter just means smarter.

    • majkinetor 1 hour ago
      There is every reason to think this. Its about available quality data. The data that was easiest to fetch was already there, then we got some more data by asking experts to create datasets for post training. Once this is over, we will get to the outer world that didn't get to hoard it for bots to take it. This will certainly change. Put on a smart glasses and record what you do to fix a pipe. In 3-5 years, rinse and repeat.
      • RealityVoid 1 hour ago
        Maybe, I'm at the point I don't know what to think anymore. But somehow, this feels still non-human. Humans don't need millions of hours of training and millions of samples to learn how do to something. We have proof that systems can deal with low-shot training. So why can't these? My point is when we have human level learning ability, then we're moving, if we expect expansion of capabilities to come from data alone, I'm prepared for disappointment.
        • sebastiansm7 46 minutes ago
          To form an human to do frontier research requires decades since birthdate
    • SamPatt 24 minutes ago
      Being good (or superhuman) at software dev and maths shouldn't be treated as though it's equivalent to other domains.

      Unlocking mastery of those domains would lead to incredible gains in all others. Whereas a fundamental advancement in many other domains doesn't spread as readily.

    • mindwok 58 minutes ago
      Just like the nerds thought about computers, or the internet, or video games, or smartphones, or crypto, or... whatever.

      The most passionate, intelligent people are usually a decent indicator of where culture is going to go.

      • dimbletimbers 35 minutes ago
        The most passionate, intelligent people 1. aren’t spending a lot of time on smartphones, video games and crypto and 2. aren’t as much smarter or better at predicting the future as the barely above average people who mistakenly count themselves in their ranks like to think.
      • aleph_minus_one 42 minutes ago
        > The most passionate, intelligent people are usually a decent indicator of where culture is going to go.

        Many very smart people who I know are quite skeptical of LLMs and the hype around them. In my observation/echo chamber, the people who are very into LLMs are rather slick, career-minded people who love to present themselves as trendsetters for the "next fancy thing".

      • throwawayy6767 16 minutes ago
        Most of those technologies were controversial in their inception among the nerds of the time. Remember Eternal September? Linux/Windows people making fun of Apple hipster fans? Mainframe engineers deriding PCs are unserious and video games as childish?

        >The most passionate, intelligent people are usually a decent indicator of where culture is going to go.

        That's a naive view of cultural determinism. The software landscape would have been very different if a handful of lawsuits had gone one way or another. Hell, even Unix and Linux were hobby projects that accidentally got big.

        Also crypto is a grift lol, one of those is definitely not like the others

      • grtt1 36 minutes ago
        Steve jobs wasn’t a nerd - he said he’s a hippie on the lost interview. He was also the first to bring to market colour screens, typography etc. He also defined the Internet as the defining ‘social moment’ way before anyone else.

        So you’re wrong pal, sorry.

    • aleph_minus_one 1 hour ago
      > Nerds (software dev, maths etc) see how good LLMs are at things they care about and assume they will be broadly applicable in future.

      I don't think this explanation is sufficient:

      For example, many nerds care about 3D printing. On the other hand, my experiments (and the experiments of many nerds who I know) to let LLMs create files for 3D printing lead to horrendously lacking results.

      Or many nerds care about linguistics or puns. Whenever a new LLM comes out to which I have easy access, I do the test, and let it explain some specific German jokes to me that are based on convoluted puns in the German grammar. Until now, no LLM that I had access to could give satisfying explanations of the puns on which these jokes are based.

      Or even for coding (many nerds do care about elegant, sophisticated code): the LLMs that I could test were already overchallenged with the following task: I had written some code that is a very "artisanal", "clever" improvement of an algorithm over the version that one would find in a textbook. The task for the LLM was simply to write some code comments/documentation about the mathematical ideas upon which my improvement over the textbook version of the algorithm is based. It wasn't capable to do this. On the other hand, for a junior programmer, I would in such a situation expect that he goes through every single line and thinks through the algorithmic ideas so that he can learn from them.

      ---

      Thus: even for "nerdy" topics, it is very easy to find tasks where LLMs still suck. So, my hypothesis about the nerds that you mention in your post is that these nerds are rather people who want to believe (with religious fervor) that the current LLMs are exceptionally good instead of just looking into a slightly different direction than where the tech billionaires want the users of LLMs to look at.

      • grtt1 34 minutes ago
        You’re having a hard time distinguishing nerds and hippies.

        Hippies live at the center of technology and humanities.

        Nerds sit purely in technology.

        • aleph_minus_one 25 minutes ago
          I honestly don't get your point: perhaps your argument is based on some US-specific cultural reference that I am not aware of since I don't live in the USA.
      • nullstyle 13 minutes ago
        > For example, many nerds care about 3D printing. On the other hand, my experiments (and the experiments of many nerds who I know) to let LLMs create files for 3D printing lead to horrendously lacking results.

        Try again; astra has been driving fusion for me. I took a screenshot of a self watering cat grass bin on makerworld — not even the stl, a screen cap of one of the photos attached to the design — and it made me a parameterized version in one shot.

        Its initial design was less than ideal for printability and it reworked the design correctly when asked to consider what it originally produced through the lens of printability.

        Autodesk bundled an mcp server with fusion in the last few weeks and i expect things to get even better going forward.

        • aleph_minus_one 0 minutes ago
          > Try again; astra has been driving fusion for me. I took a screenshot of a self watering cat grass bin on makerworld — not even the stl, a screen cap of one of the photos attached to the design — and it made me a parameterized version in one shot.

          I have no access to Astra, but with GPT 6.1 Sol, the results with respect to attempting to create STL, OBJ or even OpenSCAD files were really bad.

    • grtt1 37 minutes ago
      Correct spot on
    • bigstrat2003 1 hour ago
      > Nerds (software dev, maths etc) see how good LLMs are at things they care about and assume they will be broadly applicable in future.

      It's worse than that. I am one of those nerds (a programmer), and I see that LLMs are shit at the things I care about. I have zero reason to believe that they will be good at other things either.

  • comeonbro 1 hour ago
    I would propose another mechanism: even the free-tier models have already completely saturated what most people are capable of appreciating.
    • oh_my_goodness 57 minutes ago
      Google's free AI responses are moving the bar on what I expect from search. But maybe not in the direction you think.
    • gammarator 1 hour ago
      Or maybe needing.
    • DebtDeflation 1 hour ago
      Honestly, outside of coding tasks, the AI Summary at the top of every Google search is adequate for 99% of what I need and I hardly even use ChatGPT any more.
      • throwuxiytayq 1 hour ago
        God damn. I respect you for being honest, but… god damn. That’s a pretty low bar.
        • Jare 49 minutes ago
          It's just a google search but you can explain what you're searching for. AI summary may contain misleading details if you read it and take it literally, but 90% of the time it gives me pointer for the thing I need and I don't have to scroll through results that often could be just as misleading if not more.
        • antonvs 1 hour ago
          Just in the last few days I had an example where the Google Search version of Gemini had a much better and more comprehensive answer, by far, than ChatGPT. It had to do with Amazon’s use of Data Matrix barcodes on shopping bags. Gemini was able to fully describe how they’re used in Amazon’s logistics chain. ChatGPT essentially said it didn’t know. Your preconceptions may not be accurate.
          • causal 17 minutes ago
            People are really slow to update their priors after bad experience with an early version
          • throwuxiytayq 43 minutes ago
            They might be, but still, I’ll be difficult to convince that Google’s shitty-ass 10B params instant answer models are good at anything other than cookie recipes and producing farts real fast for real cheap.

            For what it’s worth, any time I see that feature, it’s all farts.

  • theturtletalks 1 hour ago
    Claude Code accelerated this divide. Programmers using Claude Code realized that giving LLMs access to a terminal made them feel 100 times smarter. Access to Bash, CLIs, and the ability to visit websites without the many restrictions made them powerful.

    Meanwhile, the average user was still using ChatGPT, which probably felt like it had plateaued over the last few releases. The real power of these models became apparent when paired with a terminal. Agents like Muse and OpenClaw are now attempting to bring that same Claude Code-like power to everyday users through a simple chat UI.

  • danpalmer 53 minutes ago
    2015: If we solved taxis it would transform the world completely

    2020: If we solved currency it would transform the world completely

    2021: If we solved delivering small items within 30 minutes it would transform the world completely

    2022: If we solved being able to sell jpegs it would transform the world completely

    2024: If we solved the programming part of software engineer's jobs it would transform the world completely

    • misiti3780 10 minutes ago
      Couldnt have said it better myself.
  • usernomdeguerre 11 minutes ago
    I feel like if the gap were so monumental that he could just point to a monument of it, rather than hand-waving at "the vertical professionals".

    Alternative interpretation of the numbers; there's only a subset of productive work that LLM's are 'capable' of doing at this point. That is, work that's verifiable by a machine or doesn't require verification. Thus, the cleavage points are along that axis and-- naturally-- most human beings don't have that kind of work.

  • davnicwil 7 minutes ago
    At a certain point the offhand hints at exponentials in the narratives are going to have to come into contact with reality.

    There's been this continual allusion to ever-narrowing circles of influence and ever-accelerating capabilities behind the curtain, but in fact things do keep coming out from behind said curtain and while their capability is indeed advancing it doesn't really seem at all like it's accelerating, even, let alone vertical.

    Also at what point do we raise the fact that despite the ever narrowing circles of access, nobody within those circles seems to be able to actually do anything with it that's ostensibly what it's sold to do? All they seem to be able to do frankly is build more models and harnesses and tools designed to consume more tokens from the models they build, but not actual directly useful new or innovative software that generates what could be very real revenue.

    Eventually when you keep talking about exponentials but not that much seems to be demonstrably changing, the narrative is going to run out of steam.

  • AvAn12 1 hour ago
    Fair assessment. Maybe the messaging should focus on “these are great accelerators for software developers” rather than “AI will change everything for everyone everywhere…” It is understandable that non-technical folks are kind of underwhelmed - not due to lack of understanding so much as lack of a tangible need. Not everyone needs an electron microscope or gas chromatograph…
  • Kim_Bruning 51 minutes ago
    I think we can blame the accelerating capabilities of the models somewhat.

    The Opus 4.5 model I was so impressed by last Christmas is out of date now.

    I can totally imagine someone going "I looked at this stuff a year ago and it couldn't even start a project properly" , and they'd be somewhat right.

  • m101 1 hour ago
    My interpretation of this is something like: if LLMs are to be mega useful token counts need to increase by many orders of magnitude -> broad adoption (and spending) would require token costs to drop by many orders of magnitude -> before the common folk get mega useful tools existing GPUs will be worthless
  • truthbe 35 minutes ago
    I must be the biggest failure, because I understand what AI is capable of, know how to use it, but have no idea how to profit from it without selling my time.
  • rbehrends 59 minutes ago
    > Somewhere around 20M people (0.2%) see first-hand that large, complex projects that used to take them weeks/months can now be completed by agents with a prompt.

    Really? That's not my experience. In fact, I've experimented quite a bit with spec-driven design where the LLM does not just get a prompt, but an (informal) spec, often with design considerations included. And yet, regardless of the power of the model, I've never seen them autonomously deliver what I'd consider a finished product.

    There are always edge cases that are not properly taken into account, architectural oopsies, performance issues, duplicated code, and other problems, that then need to be identified and fixed, unless it's throwaway software (e.g. a one-off script or some quick experiment).

    This is not to say that agents aren't extremely powerful (as I think they are, they do often blow my mind) but the prompt-and-forget approach is IME not a productive use of them for software that is meant to last. This may change in the future, but for now, fully or largely autonomous agentic work does not seem to be the road to quality.

    • alsodumb 38 minutes ago
      I feel like a lot of issues you mention will eventually point to an ambiguous spec. I’ve had great success with interview style spec generation where I ask agent to grill me on any and all parts of the spec/plan that’s ambiguous and then note my responses. It takes like 10 mins but the end product is always exactly what I expect and most of the things you mention get resolved in this interview phase.
      • rbehrends 11 minutes ago
        All informal specs are of course to some extent ambiguous. (I should mention that I have a background in formal methods and have also worked with formal specification tools.) And what you suggest is a process that I have used also, and it will cover some obvious gaps, but not all, at least for non-toy projects.

        If you want to go all the way, a fully formalized specification is not less work than writing the software, though the potential promise of LLMs for formal methods is interesting.

        But for any informal specs, no matter how detailed, while LLMs can be very capable when it comes to avoiding surface errors when translating them into code (e.g. they're more likely than human programmers to write code that compiles on first try), they still don't have a mental model, and most of the problems I've seen are IMO results of that and they increase with the scale of the project.

  • dofm 21 minutes ago
    Company spokesman keen for you to know that unannounced future products are definitely worth absurd valuation.
  • anukin 1 hour ago
    Tbh building an agent swarm and the coordination layer is not exactly frontier level. They don’t achieve any meaningful outcome rather than producing pr puff pieces. Hacking huggingface and Australian govt etc is very much possible with a team of humans and agents and does not need agent swarms. The cost is also lower.
  • beloch 20 minutes ago
    "So this is the weirdness of the moment. The general public has mostly not interacted with these systems. When they have, it looks like a derpy chatbot. The majority of professionals still see only a modest uplift. And a small sliver of professionals are experiencing the vertigo of the curve going vertical. And it is all happening at the same time."

    -------------

    It comes down to the difference between programmers and non-programmers. If you're not coding, AI literally is a derpy chatbot that's made it harder to get search results, raised your power bill, destroyed property values in a neighborhood near you, is propping up the economy but also creating a huge bubble that could pop and destroy a lot of retirement savings, etc.. At best, it's a nuisance, but it could be a lot worse than that. The AI evangelists need to understand this, because the average person is starting to pay more attention to the downsides while still not seeing much upside.

    This gap between the people seeing the curve go vertical and those who are simply growing more and more annoyed is a real problem... mainly for the people seeing the curve go vertical. If the people who are merely annoyed right now don't see some of the promised benefits soon, or if any of the negatives get worse (e.g. massive increases to greenhouse emissions), they might get angry. You probably wouldn't like having the overwhelming majority of the public angry at you. AI evangelists need to start doing better things for the majority of people or they need to address the negative consequences they're inflicting much more aggressively.

  • socializer 1 hour ago
    I think it's a weird take because it implies that the 6 billion people he's talking about actually have some interest in knowing about the capabilities of LLMs to solve frontier math? This is simply not something they care about or can evaluate. There are maybe several thousand people in the world who have some (abstract and barely-monetizable) use for this information, plus probably another 100,000 who don't understand any of the math, but like to cheer on.

    We have already reached "peak LLM" in terms of what normal people realistically need it to know or reason about. In fact, I'd say we reached that point about 1.5-2 years ago. There are two other barriers that remain unsolved:

    1. They're less dependable than humans and can't be meaningfully punished or forced to make up for mistakes, so you can't really replace humans with them, not without having a human babysit.

    2. Most people don't really have a special need for an LLM in their life. They may like that it answers questions or helps you polish a resume or, I guess the labs' favorite, helps you make restaurant reservations. But this sure isn't worth $200/mo for most people. Probably not even worth $5/mo.

    It'd be kinda funny if we create superhuman AGI and then no one has any real use for it, perhaps except for military murder-bots. There's always market for that.

    • HolyLampshade 1 hour ago
      I know I’m bordering on repetitive and strongly negative, but the combination of the two barriers you mentioned has all but halted all of my interaction with agents.

      Even ones that cannot be easily dismissed (like the agent at the top of Google search) are often subtly wrong in a way that requires significantly more effort on my end to parse the loquacious output to determine where the inconsistencies are.

      Look, can they be useful to generate the html or jscript for humanity’s 9 billionth iteration of a web form? Sure. But man, I wish sanity had reigned and people had applied ML models in general to more substantive and beneficial projects.

      (Before anyone chimes in, I’m familiar with implementations of ML applied to esoteric domains; but by their very nature these don’t get all the news cycles, or hiring, or any of the other absolute insanity that the domain seems to contain)

    • iugtmkbdfil834 1 hour ago
      This is by the weirdest possible opinion on llms as a whole. Not the AI booster club, not the AI doomer, not the undecided, but somehow "I can't think of anything to use it for." It is like watching the beginning of the internet in 90s and saying there is nothing to browse. The whole concept is just weird to me. At least the first 3, I can conceptually understand. It is admittedly hard for me to understand someone saying 'i tried this and best it can do is summarize emails'; I am not saying it is not true ( my boss said that exact thing ), but it is hard for me to reconcile with my view of the world.
      • nasmorn 55 minutes ago
        Also seems weird to me. Claude cowork could just do my monthly tax receipt collection and compare with my account better than the bookkeeper I pay for it. That is a very cross sectional problem
        • SpicyLemonZest 23 minutes ago
          Most people rarely have these kind of administrative tasks in the first place, and when they do they often run into GP's barrier 1. I can't recommend that my extended family have Claude do their taxes, even though I work with it on my own taxes, because they might ask for something that produces bad results and can't evaluate the output to double check. ("Hey Claude, I want a higher refund, can you tell me what to put on the forms so I get at least $3000?")
  • variety8675 1 hour ago
    A more cynical view is, trust us we've got really good stuff internally you're not allowed to see, but give us more money
  • skippyboxedhero 1 hour ago
    Text generation is not the bottleneck. Does everyone work for Accenture and TCS?
  • dist-epoch 6 minutes ago
    > For intelligence agencies, and for corporations like Microsoft, Google, Apple, Facebook, Amazon – giving you convenience and security is likely not their primary goal. I think they have an entirely ulterior motive. They can use passkeys to prevent you from running unapproved software.

    > Remote attestation can allow them to verify that your device is not capable of, say, sending and receiving encrypted messages, or running a competitor’s app

    How does this make any sense? Google can already prevent any app from running on an Android if they really wanted to, they don't need pass keys for this.

    Google doesn't need attestation to block Signal. What is this non-sense? Or will a website block phones capable of running Signal? That's all phones. And anyway, has nothing to do with pass keys.

  • tripleee 1 hour ago
    He's intentionally forgoing all nuance in order to make this sound dramatic

    > Somewhere around 20M people (0.2%) see first-hand that large, complex projects that used to take them weeks/months can now be completed by agents with a prompt.

    No, they can't, at least not any semblance of quality. The cases we're seeing where this does kinda work is in ports and translation where all the rules are already documented in the best specification language possible with a way for the LLM to verify itself: code. We saw this close to a year ago now with Cloudflare and NextJS

    > The impact scales with ambition, problem size, and horizon. A question with a paragraph answer barely stresses the system. You need a reservoir of big, difficult problems that you really care about

    These are operating on different capabilities - AI's ability to answer informational queries as a chatbot frankly sucks and can't be trusted without verifying it. I run up against this every day. A problem with a verifiable answer on the other hand it's very good at solving. He knows this (his next paragraph) but he's putting them on the same scale of "stressing the system" to attempt to add proof to his introductory claim

    And then there's the completely unverifiable scare that there are internal frontier models way beyond anything we've seen "swarms of thousands of agents collaborating over weeks on software mega projects: minting zero days, running cyber attacks and defenses at machine speeds, discovering new science, advancing the frontier of mathematics"

    I dunno. I haven't been able to set up OpenAIs remote codex connection, their shit is buggy as hell and the web UI keeps crashing and making messages disappear. Is this what their internal superhuman "Things that would have taken top professionals in the industry years of work" looks like? Granted Claude has been really smooth, but still..

    • zer00eyz 10 minutes ago
      > where all the rules are already documented

      Defined testability.

      > swarms of thousands of agents collaborating over weeks on software mega projects

      If you have defined testability, running a single agent, in a loop, over a goal will achieve the same results.

      > minting zero days, running cyber attacks ... Granted Claude has been really smooth, but still

      The old adage about monkeys writing Shakespeare applies here. Agents are just better monkeys and if the goal has a well defined outcome it's fairly easy to "Brute force" the way to an answer.

    • jfrbfbreudh 1 hour ago
      Congrats, you’ve discovered that you are not part of this group.
      • causal 12 minutes ago
        Yeah I suspect Karpathy is too generous. I’m finding a lot of devs with access to the best models are not effective with them.

        So I think there is are growing gaps even within the described groups.

      • bigstrat2003 1 hour ago
        That group does not exist. It's all hype, zero substance.
        • jfrbfbreudh 43 minutes ago
          This just goes to further show the divide.

          You’re trying to tell me it’s all hype when I have (depending on the day) a top 50 Productivity app in the iOS App Store for which I’ve written and read exactly one line of code (the version number).

          https://imgur.com/a/ua2lbxb

          • Jtarii 8 minutes ago
            It is comical that you think a good bar for quality is the fucking appstore.

            What even is this comment. Like, oh wow, an iOS app, AI hype=real.

            ???

          • SamPatt 12 minutes ago
            Well done.

            It's incredible to see AI denialists still active in late 2026.

        • SamPatt 13 minutes ago
          It's insane how confidently you can make a claim like this in a forum where thousands (tens of thousands?) of people are using these tools daily to build useful software for themselves and others.

          How can you possibly believe it's all hype? Is everyone just lying? Are we all paying $200 per month because... why?

        • tripleee 48 minutes ago
          I don't think AI is all hype at all. Opus 5.5 is really impressive. This twitter (x? whatever) post is infuriatingly hypey and just written for virality, though
  • wg0 1 hour ago
    That's fine but these companies having billions in funding by now should have rewritten chalk and other terminal libraries in Rust while bundling the whole thing in Rust as a single static binary. I'm talking about horrible Claude Code and friends.

    And then using emgui they could expose full fidelity IDE+Agent workflow interface (like DeepSeek Harness and similar) that's also exposed as MCP to be operated by another agent and JSON RPC for over the network but no. Skill issue?

    PS: Even full fidelity Photoshop is possible with such high fidelity UI that has CLI, JSON API and MCP server all in the same single binary < 90 MB. For reference, see PhotoCraft or VectorCraft or WordCraft or PdfCraft.

  • skybrian 1 hour ago
    > see first-hand that large, complex projects that used to take them weeks/months can now be completed by agents with a prompt

    Really? I start with a conversation for maybe 5 turns or so, where I ask it what would need to change, what the API might be, any database schema changes, URL schemes, and so on, and finally ask it to break it down into commits. Then I let it go, implementing 3-10 commits at a time via subagents. It usually gets the UI somewhat wrong, so there are followups to fix it. This is with Sol and Luna subagents.

    Is that what other people see?

    • einsteinx2 22 minutes ago
      I use Fable and Opus exclusively and have the exact same experience.
    • sumedh 1 hour ago
      Try using Astra, Fable, Opus.
      • derwiki 59 minutes ago
        Sol 6.1 is really good too though
  • hacker_88 46 minutes ago
    We have spent the whole evolution understanding the Real Intelligence RI
  • prometheus1992 56 minutes ago
    so what was karpathy expecting?? everyone will give up on their life and start feeding the monster in the screen? that's literally what the tiny sliver is doing; they are food for the AI model.
  • j45 24 minutes ago
    This has existed since the start of LLMs.
  • hn_throwaway_99 1 hour ago
    The fundamental question I have that honestly I haven't been able to find any answers for: With all the talk of LLM-based AI capabilities "going vertical", what evidence is there (for or against) that LLM-based approaches won't eventually "hit a wall", that is find some aspect of intelligence where humans will still have primacy, and no amount of scaling will change that.

    E.g. some prominent folks definitely think LLM-based approaches will hit a wall, perhaps LeCun most notably. I also read a guest post on Terry Tao's site (which I really liked) that argues that, for all the very impressive recent AI results, they still operate within the "convex hull" of their training data: https://terrytao.wordpress.com/2026/09/13/happy-those-able-t...

    I'm just curious if there is any actual data or evidence that takeoff (i.e. RSI, "the singularity", whatever you want to call it) is inevitable with current approaches.

    • pelican0 3 minutes ago
      I don't think there is evidence in either direction (takeoff impossible vs inevitable).

      But the "convex hull" idea always came across as wishful thinking and hand-wavy pseudo-science to me.

      Here's a rebuttal [0] of it from Tim Gowers (uses another math metaphor, but still clear):

      >...the popularity of the idea that AI "merely explores the convex hull" mystifies me. Why wouldn't it be more like the subgroup generated by existing mathematical knowledge? (That's also a metaphor -- I'm not saying there's a group structure.)

      >The convex hull metaphor suggests that AI isn't having radically new ideas and will eventually run out of things to do, whereas the group metaphor suggests the possibility of building up in small steps to something well beyond what you started with, with no obvious limit.

      [0] https://twitter.com/wtgowers/status/2108090920864268679

    • throwuxiytayq 51 minutes ago
      It’s anyone’s guess.

      But when you squint your eyes, LLMs are just a least-effort manifestation of the wider neural net direction we’ve been chipping away at for decades. I don’t give a fuck whether LLMs are here to stay; they’re proof that scaling neural nets gets you pretty far.

      The real question is: will neural nets hit a wall? As a human, I certainly hope so. But I’d bet that they won’t. I’d bet that we’re igniting the ~~atmosphere~~ universe in slow motion.

  • freecodeio 1 hour ago
    I don't understand how "swarms of thousands of agents collaborating over weeks on software mega projects" works with the current context limits and at this point I'm too afraid to ask cause I'm afraid an AI bro is gonna punch me.
    • riffraff 1 hour ago
      You can have agents make a plan and then other agents spec tickets and yet more agents do implementation and yet more do reviews and QA and whatever. It's possible.

      Is it effective? I'm not sure, cause if it was I don't see why OpenAI & co are not releasing such a project instead of demos.

    • dude250711 1 hour ago
      Your product needs to be either the swarm itself or the PR about the swarm - not the outcome.
  • ashleyn 1 hour ago
    >Meanwhile, human review and comprehension are starting to fall behind. For example, people are still involved in the "archeology" of the OpenAI-HF incident from many months ago. Mathematicians may be poring over the 722 manuscripts on frontier mathematics for a while.

    Amid all the discussion of sigmoid curves, and where the "LLM wall" will materialise, I think few people would have predicted that the real wall in LLMs would end up being humans' capacity to verify the output.

    What I fear is that people simply eschew human review altogether, considering we're talking about the industry that came up with the "move fast and break things" credo. Human review of LLM-produced code where I work is already a farce, and we're not special enough to be one of Karpathy's 5,000. I do my best to manually review anything that's my responsibility, but I'm literally one of very few people left working on my team, so in practice what happens is I submit PRs that are at best glossed over by completely unrelated teams for security, malware/prompt injection, and other serious concerns. Quality insofar as vetting others' code has completely gone out the window and it shows in the number of bug reports that come back, often themselves written in Claudease. This is all on top of everyone cynically phoning it in in the first place, due to the omnipresent sword of Damocles that is additional AI-driven layoffs.

    Worse yet all the incentives point to this being the most economically viable thing individual companies can do. I think it goes without saying some type of regulation here is urgently needed, and that an unexpected cause of an AI bubble pop may end up being that humans simply aren't able to keep up with the pace of the output - leading either to precautionary plateauing of capability, or major liability risks related to a decline in quality.

    • michaelchisari 1 hour ago
      | few people would have predicted that the real wall in LLMs would end up being humans' capacity to verify the output

      That was the dominant concern in the circles I’m in, so it’s worrisome it’s being treated as rare.

      • JBits 1 hour ago
        I would question the narrative that humans lack the capacity to verify the output and would instead argue the people lack the incentive to verify the output.

        The response of many mathematicians to the recent dump is a good example: verifying these proofs amounts to unpaid labour for OpenAI and wastes time that could be spent doing publishable work which ultimately results in money or personal success. The slop factor also compounds the work required to verify the output considerably.

        For mathematicians, programmers or anyone, if the work required to deal with slop passes the limit, it is no longer in their own self interest to use LLMs. The expectation that people will use LLMs for the betterment of humanity against their own financial interest is baffling.

      • skydhash 1 hour ago
        Humans are not immortal and cannot spend all their time into review (especially unpaid). Even today, there’s so much knowledge around that you have to be specialist of a narrow domain to get to the frontier. Even in computing which is just approaching a century of existence.
  • chevman 1 hour ago
    I mean in late 2020/early 2021, Altman and others were saying the end of work was 6 months out.

    That clearly didn't happen :)

    • Legend2440 1 hour ago
      Altman didn't say that. In March 2021, he did make some predictions, but they were much vaguer and farther out:

      > In the next decade, they will do assembly-line work and maybe even become companions. And in the decades after that, they will do almost everything, including making new scientific discoveries that will expand our concept of “everything.”

      https://moores.samaltman.com/

      • karahime 28 minutes ago
        To add onto this, I feel like this has been happening a lot, where critics project backwards imagined claims or go nutpicking for random comments online to make something sound absurd and like everything is behind schedule and nothing is materializing, and then you can see other people (not necessarily the same people, the point isn't "look at these hypocrites") saying that it's all moving too fast. Both of these can't both be true in their strongest forms. I think what they have in common is that because they're relative measures to an unstated baseline, they're free criticisms you can throw at any development whatsoever, regardless of what is or is not turning up.
  • dude250711 1 hour ago
    Am I the only one who thinks "yes, it is currently overhyped, but no, it will get there soon"?
  • notjes 1 hour ago
    [flagged]
  • mccoyb 2 hours ago
    If the software coming out of OpenAI and Anthropic is what we have to judge, I wonder about the 5000 ...

    Let's say, for the sake of argument, that the models are some multiplicative factor better on the inside.

    Doesn't that mean the demos should work?

    • j2kun 1 hour ago
      Unfortunately, marketing, hype, and venture capital overshadows any serious public discussion of capabilities.
      • AndrewKemendo 1 hour ago
        Be the change you want to see in the world: Attend or host an AGI society event to have that conversation:

        agi-society.org

    • spiderice 1 hour ago
      I'm confused.. are you suggesting that Claude Code / Codex don't work? Because if you're still saying that in October 2026, it's a you problem. You're doing something wrong.
      • mccoyb 1 hour ago
        No, I’m talking about the recent DevDay.

        Also, yes there are still bugs in Claude Code. I experience them nearly everyday.

        It is markedly better than early days, but still not the best harness.

        The best software written with agents seems to come from people outside of the labs (see pi, for instance — or all of cloudflare’s recent work)

        Which makes me question either the model, or the holders …

        • nullpoint420 1 hour ago
          Cloudflare is where you lost me. I don't know anyone actually using them other than for their proxy, DNS servers, or DDOS protection.
          • mccoyb 1 hour ago
            See, for instance: https://news.ycombinator.com/item?id=49182996 by Kenton Varda

            Also, not mentioned in my post:

            - Mitchell Hashimoto

            - Prime Intellect (and all their agent experiments)

            - Geoff Huntley (see Jiti, for instance)

            (many more)

            There's a ton of interesting software being developed with these models by people outside the in group, but I find most of the software from these big guys to be ... bland. Buggy copies of copies.

            • nullpoint420 1 hour ago
              I don't see the Cloudflare stuff as groundbreaking, let alone people building companies around it. "Workers" are the wrong primitive, IMO. What if my code isn't in Javascript? They sandboxed the wrong part of the machine.

              I have a version of "Cloudflare OS" running in production, using Temporal as the orchestration and MicroVMs for the sandboxing.

              Sorry, I get your point but personally I don't get the hype around them.

      • lillesvin 1 hour ago
        I don't use either myself but I'm in a position where I get to see what it produces for other people. Here's a recent (2-3 months old) example: A zsh script that immediately invoked a Python interpreter that immediately invoked a shell command and iterated over each line of output looking for one of a few strings to match... So, `<cmd> | grep "(stringA|stringB|stringC)"`, but a whole lot dumber.

        Thanks, but no thanks, I don't need that kind of code in anything I'm working with.

      • well_ackshually 1 hour ago
        Claude Code still doesn't have a working scroll back buffer in their new renderer, it regularly flips its shit and mixes different pieces of history.

        Claude Code still can't get reasonable performance without writing a "game renderer" (that doesn't work)

        Claude Code is still written in JavaScript, eating hundreds of megabytes to make a shitty TUI whose literal sole role is to send API calls.

        Claude Code is software made by amateurs.

      • plorkyeran 1 hour ago
        Claude code is an incredibly buggy mess. I have never used any other TUI that regularly has rendering errors or that is anywhere as laggy as Claude.
    • LastTrain 1 hour ago
      It’s like the aliens paradox. If AI can build killer software already, where is it?
      • equinumerous 1 hour ago
        Couldn't agree more. I find a new bug in the VSCode Codex extension every day... quantity != quality!
        • agentdev001 1 hour ago
          I have a feeling that the team working on the VSCode Codex extension is one guy, who begrudgingly takes down a couple tickets every other week.
      • shermantanktop 1 hour ago
        One solution to the aliens paradox is that they are so advanced they can hide from us. Maybe the killer software is kept inside the labs? I doubt it.
  • AdeptusAquinas 1 hour ago
    "running cyber attacks and defenses at machine speeds"; worth noting that LLMs (even the vaunted frontier models) are exponentially slower at cyberattacks or defenses than your average WannaCry or Splunk automation from decades ago. Its this sort of delusion world the AI bros live in that is part of the reason there is a disconnect between what they think should be happening and what actually is.
  • CapitalistCartr 30 minutes ago
    [dead]
  • Apollorider 22 minutes ago
    [dead]
  • kydanet 2 hours ago
    [flagged]
  • rolosa 1 hour ago
    [flagged]
  • 8484848484 2 hours ago
    [dead]
  • weinzierl 2 hours ago
    [flagged]
    • atmavatar 1 hour ago
      > a tiny group is watching the curve go vertical

      Caveat: that same tiny group is employed by the AI vendors, meaning it's in their financial best interest to make it sound like the curve is going vertical.

      • pianopatrick 1 hour ago
        Also, part of "seeing the curve go vertical" is that the projects that show those results are rather expensive. If you work at one of the AI companies or certain well funded customer companies then you can pay for the tokens. But, like it would make no sense for me personally to spin up a bunch of agents to try to write a browser or solve math problems.
      • Arkhaine_kupo 1 hour ago
        Down is a perfectly valid direction for a vertical line when not given a ± in the vector.

        Considering the investment in AI, the lack of moat, and the increased inability of any of the big players to come even close to profitability (with OpenAI already breaking the "ads" emergency glass option)... perhaps he meant a tiny group is already seeing the line crater

      • majkinetor 1 hour ago
        It can be both and it is.
      • njovin 1 hour ago
        Another caveat: many of that same group seem to have a shared delusion that they’re birthing a super intelligence, and those are the same ones claiming the vertical curve.
        • sumedh 1 hour ago
          > and those are the same ones claiming the vertical curve.

          Do you still write code by hand?

      • kmac_ 1 hour ago
        I work for a typical software product company, and along with most of my colleagues, I clearly see that the curve is so steep that our software development process has already changed tremendously and will be different next year, and probably completely different in the following years. The revolution is real, undeniable, and the old days are gone. Some companies adapt to changes more slowly, some faster. LLMs, agents and harnesses are just a part of the bigger picture.
      • nullpoint420 1 hour ago
        I hate to say it but this is cope. I believed this in the past but it's over.

        AI models can dismantle billion dollar industries. They can reverse engineer Adobe and Microsoft products that once were their moats and titans of their industry.

        Why do you think they'd need to lie?

        • riffraff 1 hour ago
          Because if they can, why haven't they?

          LLMs are great and will get better but we have not seen an open source office suite built by LLMs come online yet, and even if there was one I'm pretty sure few would adopt it, cause that's not the only moat, LibreOffice has been around for decades.

    • msy 1 hour ago
      The curve go vertical for what, precisely? For all the chest-puffing and ominous and cryptic comments from 'insiders' about these incredible capabilities every time they put something public it turns out to be a pale shadow of what was trumpeted. These are powerful useful tools but the quasi-cultish behaviour around them is getting old.
    • albatross79 1 hour ago
      Congratulations, you've parroted something said by someone else.
    • lifeisloving 2 hours ago
      I use models all day everyday, have unlimited access to all models. The curve is not going "verticle". I have all the workflows and meta agentic tooling, im not holding it wrong. Its bad, not everything is a 20th percentile problem.

      There is in fact no indication of this, not evem the precious benchmaxxed benchmarks ya'll love to reference.

      There is however a exponential curve of slop, and an ever increasing number of people who's minds are completely captured by these things.

      • tkz1312 1 hour ago
        As someone who has done software verification professionally for many years the last 6 months or so have looked extremely vertical. The robots are better proof authors than I probably ever could be even if I dedicated the rest of my days to the practice, and projects that once would have taken months now take a day or two.
        • lifeisloving 1 hour ago
          Then you'd know well that 2hr of LLM code generation can easily be about 4-8hrs of review, and that review can be brutal.

          I'm not arguing that they cant write code, or write a proof. Its just not written or designed well and is absolutely brutal and soul crushing to work with. Look at these proofs they're producing also, they're millions of lines of Lean that are impossible to reason about.

          The way we're using the term 'verticle' to describe a curve means we're not being honest about this. This curve can actually be plotted, you can go look at the curve. It is not in fact 'verticle'. Each model release is climbing single digits on benchmarks it was overfit for.

          • tkz1312 1 hour ago
            You don't need to review proof code.

            In the last 9 months or so llms have gone from just another useful proof tactic (like grind or sledgehammer or sat solvers) to being so good at writing proofs that I don't even bother to try myself anymore.

        • gr_norm 1 hour ago
          I believe this, but it is also a unique case where the pitfalls of LLMs (producing weird errors that a human wouldn't) are zeroed out. Since you have a proof checker that tells you if the LLM did it right.
          • tkz1312 59 minutes ago
            I'm pretty convinced most serious software will have some kind of proof system inside within the next few years.
  • kittikitti 2 hours ago
    [flagged]