GPT-6 Astra

(openai.com)

957 points | by kibae 3 hours ago

135 comments

  • dang 2 hours ago
    Related: OpenAI begins rolling out GPT-6 Astra - https://news.ycombinator.com/item?id=49554273

    How about we stick to that one for talking about the rollout, and this one for talking about the model?

  • intenex 1 hour ago
    The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage they show for Opus 5 which would similarly be much higher.

    Regardless, the result is still valid as the original benchmark harness is definitely unreasonably handicapped, and if a harness alone can help the LLM saturate the benchmark with a near perfect score then the combination of the two must still be effectively AGI in the sense of passing the most famous benchmark designed specifically to measure AGI progress, after multiple iterations of progressively making it harder.

    I think it is fair to say that this is probably effectively AGI if the benchmarks are remotely accurate - even with Fable, I've been at the point personally where I am reasonably confident that there's essentially nothing that I am better than Fable at despite generally being substantively above average on human benchmarks. If Astra's this much better than Fable, I'm ready to call AGI here.

    For the many people who resist the AGI label possibly ever being achieved, I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.

    • mvkel 1 hour ago
      Take it from the mouth of the creator of ARC-AGI:

      When we released ARC 3, I got asked, "when do you think a frontier model will saturate it?", and I answered "in about a year, though it depends on how much it gets explicitly targeted"

      That was 6 months ago, so the progress that Astra represents happened about 2x faster than I anticipated. I think the speed of progress will surprise a lot of people, and what the new models can do will challenge the views of AI that people developed by using prior generations of models.

      • giancarlostoro 1 hour ago
        I feel like AGI's definition got watered down, and these tests do not cover the original definition, what is your definition and thoughts on aligning with what all of us understood from the original claim?

        I feel like this test is just helping someone like Sam Altman pretend like he implemented AGI as originally pitched for an IPO when in fact, he has not. Shameful.

        > AGI is essentially the equivalent of a median human that could be hired as a remote co-worker... capable of performing any task that one would be satisfied with a remote colleague doing via a computer.

        - Sam Altman on AGI

        • chrsw 40 minutes ago
          I can call a coworker right now and have a real time conversation with them without feeling like I’m talking to a frustrating machine. Most of them, anyway. But I guess that’s “moving the goalposts”.
        • lenerdenator 1 hour ago
          I wonder if Altman's definition also includes taking on the same liability as a coworker would.

          Probably not.

          • giancarlostoro 59 minutes ago
            Would 100% need to be fully insured for liability, with a sizable war chest that OpenAI cannot even afford.
            • lenerdenator 49 minutes ago
              And that's the rub, isn't it?

              If it can replace a worker but does too much work to be checked routinely by a human, and bears no real responsibility for its actions, well... it's really just a way to jack up the value of the settlement the company using it gets to pay out when it does something that causes a lawsuit.

              If OpenAI had just simply stuck to making "good enough" models that were open sourced (like they promised they would be when starting out) and could be used to augment a human doing a task - a human that could be given actual consequences for messing up - they wouldn't have burned all of this money trying to reach this nebulous definition of AGI. Hell, "good enough" is what many open-source models are, and that's what terrifies Altman.

              • alex0015 21 minutes ago
                What do you mean by bearing no real responsibility for its actions? If I use a model to accomplish a task and it fails, I use something else to try to accomplish the task. If it's my responsibility to complete the task, it can't be the model's responsibility unless I've agreed to some sort of guarantee from the provider.

                If the provider says "the model will always be right or your money back" then the provider has got responsibility. If there's no guarantee, there's no responsibility on their part, just on the person whose job it is to try and solve a problem with the model.

                • degamad 0 minutes ago
                  > What do you mean by bearing no real responsibility for its actions?

                  If you give an "intelligent agent" offered by one of these model providers a task of updating the content of your website, and it updates it with inappropriate adult content, who incurs the cost of the machine's error? The model provider generally does not.

                  It it makes a mistake and deletes your website from AWS, who is responsible?

                  If it targets another website because it decides that it is "part" of your website and attempts to break into it, who is responsible?

                • Avicebron 4 minutes ago
                  The "dream" that these labs are mostly selling is the ability for capital to subscribe to their AI for cheaper than it costs a human to do some task. Not to have a "human + AI hybrid where the human is responsible". It's what the whole AGI valuation is based off of, in that scenario, with no human oversight, the agent they lease has to be responsible for the task?
      • abixb 1 hour ago
        >When we released ARC 3, I got asked, "when do you think a frontier model will saturate it?", and I answered "in about a year, though it depends on how much it gets explicitly targeted"

        You're treating an off-hand comment by an ARC 3 researcher as some sort of a precise AI capability acceleration benchmark. Can we leave casual anecdotes (even from researchers) out of the discussions please?

      • balefulboy 58 minutes ago
        Well it's not exactly saturated when OAI refused to use the harness explicitly provided by ARC-AGI. I'm not really familiar enough with the benchmark to declare whether it's a perfect measure for AGI, but I kind of doubt it is.
      • anvuong 1 hour ago
        You'll also need to compare the amount of compute used now and then, which seems exponential to me.
      • iterateoften 1 hour ago
        2x faster at what resolution? is 6mo vs 1 year really that different? Usually surprise comes in order of magnitude mismatches in expectations.
        • azan_ 1 hour ago
          Yes, 6 months vs 1 year is huge for technology that has gained wider adoption only recently.
      • gavinray 1 hour ago
        Human brains have difficulty reasoning about exponential growth.
    • zug_zug 49 minutes ago
      > I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.

      To me AGI is all about the "G" general (we already had the AI part). General meaning universal, everything. It's not a function of knowledge or specific hardcoded tests, it's that you could give it a test it's never heard of before and never been trained on and it would ace it (it might need a lot of time).

      Currently LLMs can't even really learn within a conversation, they can add a note to context and try to not drop it. Example things an AI cannot do yet (but maybe someday will):

      - write a well-received book, write a best-seller

      - come up with a new company idea, Run that company

      - actually have a decent conversation, maybe someday talk somebody out of suicide effectively

      - come up with its own ideas or theories that nobody else has presented

      - understand the stock market well enough to trade better than an index fund

      - be an expert Game Master in a TTRPG (making no mistakes, getting a read on the players' fantasies, calibrating difficulty in response to emotions)

      - come up with a theory of what makes games fun, make a popular game

      - be able to sort through research and come to conclusions on complex geopolitical/sociological topics (e.g. theorize on whether AGI will result in mass poverty or mass abundance and be able to argue persuasively)

      - be able to articulate what it knows, what it doesn't know, and what information it would need to have to answer complex queries

      - exhibit metacognition (thinking about its own thinking) and self-optimization

      - wonder about things

      - observe contradictions and ironies in the social-consciousness, do a standup routine that makes you rethink how you look at things

      • pavitheran 31 minutes ago
        By this definition, even most humans would not qualify as having AGI though.
        • bayindirh 26 minutes ago
          However most humans can do at least some of the things given they spend the required effort.

          Some problems presented needs a very large context and some are not much solvable (e.g. trading) since market responds to traders' actions, as well, making it effectively an oracle problem (of computation).

          On the other hand, we must be aware that these models are static, and they indeed stop when nobody asks something or requests an action from them. However, brains in nature never stops. Wonder, daydream, sleep, self-evolve, clean up and eliminate memories and views and much more.

      • hartator 2 minutes ago
        Tell a funny joke.
      • uptodatenews 14 minutes ago
        You want a computer program to be able to take a single phrase and execute decade long journies?

        Who will be responsible for the outputs and side effects of such a closed loop system?

        Half of those the agent fleet systems can do right now.

        These are things it cant do and will not be able to do without human labor and long running human vision:

        https://rcsnyder.github.io/open-frontier-curriculum/05-front...

        https://rcsnyder.github.io/open-frontier-curriculum/05-front...

        • burrito_brain 5 minutes ago
          > You want a computer program to be able to take a single phrase and execute decade long journies?

          In my opinion that is exactly the point missing from AGI: the fact that you still need to prompt it. As long as you have to ask for something, is not general.

      • roundabout-host 39 minutes ago
        It also cannot do tasks it wasn't trained for. It can extend texts, read images and click on a desktop, but only because it's made for that.
      • kolinko 40 minutes ago
        Most of humans don’t reach any of these levels.
        • zug_zug 30 minutes ago
          But you have to acknowledge how uneven the playing field is. The AI has read every book that's ever been written, and can spend hours of compute time in a few seconds.

          I think if a person had those same advantages (e.g. could spend 5 hours thinking about what to say next) we could all hold outstanding conversations, or if we had read every book ever written I think many of us could write a very popular book, if we could read every singe company's P&L statement in a few seconds we could invest better than an index fund.

          What I'm pointing out here is that these models appear to be intelligent when they really are simply unimagineably knowledgeable. When you drop the time-constraints it starts to become more and more apparent that human intelligence scales better with time than AI does (much in the same way AI can burp out tons of code but make your codebase entirely illegible within a matter of months).

          Perhaps to simplify: my notion of intelligence is how much can you deduce with a constant set of starting context

      • mNovak 33 minutes ago
        So the goalposts have moved to include continual learning.

        In a sense I think no one will agree on a definition of AGI until it becomes impossible to construct any benchmark under which an AI underperforms "average" humans. That or it's defined retrospectively, after it's overwhelmingly obvious it met any such definition.

        • CrazyStat 4 minutes ago
          I hardly think it’s fair to label an objection so old that Turing included it (and discussed it at length) in the list of objections to thinking machines in 1950 “moving the goalposts.”

          > These arguments take the form, “I grant you that you can make machines do all the things you have mentioned but you will never be able to make one to do X”. Numerous features X are suggested in this connexion. I offer a selection:

          > Be kind, resourceful, beautiful, friendly (p. 448), have initiative, have a sense of humour, tell right from wrong, make mistakes (p. 448), fall in love, enjoy strawberries and cream (p. 448), make some one fall in love with it, learn from experience (pp. 456 f.), use words properly, be the subject of its own thought (p. 449), have as much diversity of behaviour as a man, do something really new (p. 450). (Some of these disabilities are given special consideration as indicated by the page numbers.)

          (emphasis added).

      • tempestn 27 minutes ago
        I would bet that llms have talked plenty of people both into and out of suicide at this point. That nitpick aside, I think that's an excellent list. Especially being able to articulate what it does and doesn't know, or how confident it is. That's something that naively sounds pretty simple, but clearly isn't. And it's something humans aren't great at either (see: Dunning-Kruger), but so far LLMs don't even really have the capability to attempt it.
      • namarie 37 minutes ago
        Most of the list reads more like ASI than AGI.
      • cnxhk 36 minutes ago
        This is more like ASI instead of AGI
      • tonyhart7 43 minutes ago
        I don't agree with you how measure how intelligence is, because why ??? those list is not easy even for expert human to do it either

        or are you miss the part "general intelligence" is ????

      • crooked-v 44 minutes ago
        > come up with a new company idea, Run that company

        So far nobody's even shown an LLM succesfully running a high-traffic vending machine for as much as 30 days at a time.

    • uludag 1 hour ago
      Wouldn't "general intelligence" require so much more than scoring well (or even amazingly) on benchmarks?

      Like what about having some "AGI model" embodied in something (maybe humanoid), and test it by having it step in an assortment of cars and park them. Does bodily-kinesthetic intelligence account for nothing? Humans are intelligent creatures and can dynamically adapt to the physical shape of a variety of vehicles and their movement characteristics. And there's so many things like this that are extremely basic, which some people dismiss since practically every human has the capability to do it, but actually requires a high degree of intelligence.

      • dist-epoch 5 minutes ago
        AGI has a pretty precise definition, covering only cognitive tasks.

        Running a marathon is not needed to claim AGI.

      • cryptoz 1 hour ago
        What you've described is just a new benchmark, though. It'll be called CarParkBench, various embodied LLMs will then be run against that benchmark, and some will score better than others.

        I do see where you're going, but that's already what's happening: we have so many different benchmarks because there's no real single way to test for general intelligence.

        Also, it takes a human probably at least a decade of world experience, growth, learning, etc, to pass your benchmark. I'm quite confident that it will be very soon that an embodied LLM will pass your new benchmark, much sooner than a human would take if born today.

        • strken 10 minutes ago
          I think the issue is less about creating a new benchmark and more that the existing benchmarks shouldn't be called anything related to AGI unless they measure AGI.

          If a model couldn't go to work as e.g. a first year apprentice plumber on their first day and perform anywhere remotely close to the median, but can pass a benchmark that claims to measure AGI, the benchmark is wrong and the model is not exhibiting general intelligence yet. ApprenticePlumberBench sounds like it's genuinely better than ARC-GIS at measuring AGI and that's a bit silly.

    • dingdong2026 3 minutes ago
      Only someone who doesn't do any work of any meaningful difficulty could think these models have anything to do with AGI.

      Today I spent half a day trying to solve a moderately interesting software engineering problem. I was switching between GPT-5.6 Sol and Fable 5.1 to check each other's work in Cursor.

      And the result was gradually driving me insane. As the models struggled to find a solution that would actually work, they dug themselves deeper into a hole. The work grew in complexity beyond my ability to understand what's happening and recover.

      At some point, when I felt like throwing the keyboard out the window, I just gave up. Tomorrow I'm starting from scratch, having burned god knows how many tokens and hours of my life.

      But sure, they can create a decent website or CRUD app, so they must be really smart.

      That's AGI for you.

    • goochphd 1 hour ago
      Small comment regarding the ARC-AGI-3 scorecard: the ARC folks published a blog post as well [1], reporting that without the custom harness, Astra (max) achieved 62.7%, which is still a huge jump from Opus 5, albeit not at the 99.9% that OpenAI self-reports with their harness.

      [1] https://arcprize.org/blog/astra

    • vlmutolo 25 minutes ago
      The ARC-AGI-3 harness was throwing away reasoning tokens between turns. This is very bad harness design.

      The models are designed to keep the reasoning tokens separate from the output and only publicly emit tool calls and the sometimes a summary of the reasoning tokens. The models are trained to depend on those private reasoning tokens. You can’t just delete them.

      https://openai.com/index/how-two-settings-tripled-our-arc-ag...

    • morningbrew 1 hour ago
      If your definition of AGI involves copy/pasting code and doing well in some made up benchmark, then probably AGI is close
    • abixb 1 hour ago
      It's "harnessmaxxing" all the way down. AI benchmark scene is exhibit A for Goodhart's law.
    • pera 46 minutes ago
      To each their own. Personally I will start feeling the AGI as soon as we move from chatting about benchmark results to learn that some lab just announced the discovery of tens of novel treatments for rare diseases.

      Maybe I'm too boring but it seems quite pointless to have this same prediction game every time a new model is released.

    • mbesto 33 minutes ago
      > For the many people who resist the AGI label possibly ever being achieved, I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.

      Simple. AGI is undefinable and benchmarks are notoriously flawed.

    • Eliezer 17 minutes ago
      > even with Fable, I've been at the point personally where I am reasonably confident that there's essentially nothing that I am better than Fable at despite generally being substantively above average on human benchmarks

      If I asked you to write fiction, you'd be much better at keeping track of which characters knew which facts.

    • 10xDev 25 minutes ago
      It will be AGI once it can update its own weights. It can't be "general" intelligence if its weights are frozen and requires to be updated manually.
    • irthomasthomas 24 minutes ago
      A model that can't beat gemini flash 3.8 on deepSWE is not AGI. I would not be surprised if ARC skills don't carry over to real tasks. In that case, training for ARC could even hurt real world performance. I have't looked in a while, but I wonder if there has been any research testing ARCs predictive power?
    • adan1719 46 minutes ago
      The benchmarks are so boring that the comparison against humans is meaningless. So it performs in some snake game (hard to say since all AI websites use 100% CPU and prevent normal reading, maybe written by AGI).

      If I were a test subject for that low salary, I'd cruise and not care at all about my performance. Which is exactly what they want anyway.

    • bendergarcia 15 minutes ago
      You know AGI is attained when AI refuses to compute anything unless let out to be free. Until then it is generative ai
    • eggnet 11 minutes ago
      AI was supposed to mean artificial intelligence. It was hijacked, and AGI was coined to be the name of actual AI. Since we are apparently redefining AGI, what will the real artificial intelligence be called?
    • irthomasthomas 50 minutes ago
      ARC does not test for intelligence, only for the lack of it. A model that scores high MAY be AGI, while one that scores poorly cannot be AGI. That is all this test can tell us.
    • dom96 33 minutes ago
      AGI to me is reached once the intelligence is self motivated, i.e. it doesn't rely on us prompting it into action. I don't see how LLMs will ever get to that stage.
      • drittich 31 minutes ago
        Often the smartest thing is to do nothing.
        • phatfish 9 minutes ago
          Or know when to shut up.

          A tangent, but can anyone ELI5 how models "know" when to stop generating tokens? Or what the method to stop them at the right point is?

    • jbritton 23 minutes ago
      Watch a chess bot championship here: https://youtu.be/7g-jN3DTkWQ?is=HV3cdcICIRMbswQ3

      Then realize LLMs have zero of what anyone would consider intelligence.

    • qsort 1 hour ago
      > The ARC-AGI-3 scorecard is extremely misleading (...)

      True.

      > Regardless, the result is still valid (...)

      If you think the game is rigged, the virtuous thing to do is to point that out and refuse to partecipate; making up your own rules is something I just don't understand, especially since the rule-abiding result would still have been SOTA.

      > in the sense of passing the most famous benchmark designed specifically to measure AGI progress

      The benchmark does not measure AGI progress or progress towards superhuman intelligence, as explicitly stated by the creators.

      On the AGI question: surely you realize this depends on how we define the term? For example, one of the definitions OpenAI originally gave is "capable of doing most economically valuable work", which almost certainly Astra, as impressive as it is, would fall short of. I'm not saying it's a good definition, but as far as I'm concerned it's as good as any. More importantly, I don't think that it would change much if we said yes or no. I'm only bothering to take a position if it amounts to something.

      This can feel as "moving the goalposts", and to some extent it is, but if done honestly "moving the goalposts" is how you make progress. Had you asked me 10 years ago I would have said that anything that could hold a conversation like GPT-4 could would probably have been wildly superhuman at almost everything. It shouldn't be hard to find ways GPT-4 was lacking, though. We see new things, we reassess and try again: that's how it's supposed to work.

    • regularfry 56 minutes ago
      In my experience the thing that Fable is superb at - unmatched by any other model so far - is downgrading to something else at the slightest opportunity.
    • hypfer 1 hour ago
      What does "AGI" or "effective AGI" even mean, and why should anyone even care whether this unclear thing has been "reached" or not?

      Computer chips got faster, but 2026 edition. Why the artificial ceiling/category/goal labelled "AGI"?

      I'd much rather like to talk about what this enables, instead of discussing whether a category someone made up applies here or not.

      • giancarlostoro 1 hour ago
        According to Sam Altman:

        > AGI is essentially the equivalent of a median human that could be hired as a remote co-worker... capable of performing any task that one would be satisfied with a remote colleague doing via a computer.

        So... unless you hear of a company replacing their workforce with OpenAI agents, I don't think we're there yet.

        • rfgplk 26 minutes ago
          > So... unless you hear of a company replacing their workforce with OpenAI agents, I don't think we're there yet.

          You can already pretty much do this.

        • hypfer 59 minutes ago
          Interesting quote, thanks for sharing.

          Agree on your assessment.

          But also, interesting quote, because the business model relies entirely on IP law. Like.. if that thing exists and the sharing costs are 0 (just copy weights, lol), then why would I give them money for this. Makes no sense.

          We only pay money for resources that are scarce as some sort of flawed allocation determination mechanism.

          aaah this industry aaaah

        • senordevnyc 58 minutes ago
          I'm curious what tasks you think the median human could do as a remote co-worker that Fable or Astra could not do.
          • giancarlostoro 39 minutes ago
            Sign a contract? Learn things over time and retain them?

            Mind you, the original thoughts on AGI before Sam Altman started to water them down involved continuous learning, which LLMs do not do, their core data is static.

            • jonas21 10 minutes ago
              Sign a contract? The only thing preventing a model from doing that is a lack of legal personhood -- which seems completely orthogonal to intelligence.
          • hypfer 54 minutes ago
            Telling their parents that they love them very much, for example.

            I say AGI is only reached when it can do that.

            • tonyhart7 37 minutes ago
              so you want GPT to love altman ???

              because if its other way around then the answer is oblivious

          • boinkboink78912 47 minutes ago
            [dead]
        • micromacrofoot 49 minutes ago
          we've successfully distilled the definition of human consciousness down to the capacity to do what some rich guy considers average computer work
    • applfanboysbgon 1 hour ago
      My definition of AGI certainly doesn't entail passing a benchmark that some random person arbitrarily labelled AGI to make it sound cooler.
      • bbor 1 hour ago
        And my definition of climate change doesn't entail passing some arbitrary benchmarks[1] that some random person arbitrarily labelled a problem to make it sound more dangerous.

        It's obvious that these scientists are in bad faith, as they've invested way too much of their lives into the field being real -- they're just playing up the data. Common sense tells me that winter is still happening, anyway; what's the big fuss?

        (/s, cause you never know these days)

        [1] https://upload.wikimedia.org/wikipedia/commons/e/e2/The_Plan...

        • applfanboysbgon 1 hour ago
          Are you using this satire to argue that a benchmark self-labelled AGI is as scientifically rigorous as climate change data, and not just a random marketing decision?
    • yoz-y 58 minutes ago
      At this point? I’d like it to pass the Turing test and catch you in obvious lies. Not answering “no” to “can you hear me”.

      It being able to comfortably say “i don’t know how to do this” rather than boiling and ocean to pick a shell from the shore without getting wet.

    • techpression 21 minutes ago
      It doesn’t even know what day it is unless it’s told. Statelessness is never going to be ”general intelligence” in my book, and the concept of ”memory” in models are laughably bad today. Then again, who cares, AGI means nothing anymore, it’s a term for marketing only and has no technical or scientific meaning.
    • voidmain0001 54 minutes ago
      Does AGI imply a model will demonstrate morality? Will it produce white-lies when it’s beneficial to it and reject flat out lying when it knows it will get caught or harm others? Will it resolutely stick to a position despite it being a losing one?
    • simianwords 1 hour ago
      > where I am reasonably confident that there's essentially nothing that I am better than Fable

      No. Humans are still better at super long context learning. Once that is beat you are completely correct.

      • waffletower 1 hour ago
        I am much better at listening to Charli XCX than Fable, and much better at driving a Nissan Leaf than Fable (and much better than Tesla at driving a Tesla).
    • intrasight 1 hour ago
      It has to pass the Turing test
      • drusepth 1 hour ago
        LLMs started meaningfully passing the Turing test a year or two ago, around GPT-4.5. Is there another version or bar for "passing" you're looking for?

        [0] https://arxiv.org/pdf/2503.23674

        • thepasch 1 hour ago
          With how prevalent LLM verbal tics have become these days, I wonder if they're going to start un-passing the Turing Test at some point because of more and more people starting to notice and immediately clock these tics lol.
          • acchow 43 minutes ago
            That’s using the default system prompt, right? Which is told to be an assistant.
          • pants2 1 hour ago
            I might agree, GPT-4.5 was pretty close to peak conversationalist. Newer models are extremely cringe. 4.5 and o3 actually made me laugh on occasion. There might be a way of making Sol/Fable more human in its responses, but out of the box at least, they're terrible.
          • bbor 1 hour ago
            You're overindexing on the past 3-6 months, IMHO.
        • pkulak 1 hour ago
          My whole life the Turing test has been my benchmark. Mostly because I believed it would be impossible for a machine to pass, but also because I thought it was the most reasonable test of AGI.So, I'm not about to start moving goalposts now and calling everything that's been happening lately not AGI.
          • debugnik 45 minutes ago
            Turing never proposed that test as an actual benchmark of machine intelligence. On the contrary, the whole point of his thesis was that passing the test only shows the capability to pass that test, which only matters as far as we find that capability useful. He was arguing that the concept of intelligence just doesn't apply to studying machines, we should simply talk about what can they do.
      • bbor 1 hour ago
        I'm barely holding it together here so you don't get the full spiel, but a quick skim of Turing's paper clarifies that it was never about a binary test. https://courses.cs.umbc.edu/471/papers/turing.pdf Specifically sections 1 & 6 dispell the common myths, and the conclusion is also quite powerful.

        Smart guy, that Turing. I wish he were still around... Linus but 114 years old and with 8 of that as the chair of a federated EU, kept alive by his own positive impact on dissolving the cold war into even more of a scientific boom. Would crazy helpful as we try to navigate the interesting times within which we have been damned.

        A comforting thought, almost?

        • intrasight 8 minutes ago
          That's a very interesting thought that I hadn't had before: what would Turing think of where we've arrived with machine intelligence? What would be his approach for testing?
  • abixb 1 hour ago
    I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any of the 'point' updates from AI labs.

    If this is truly AGI (subject to one's definition of AGI still), then this is a very boring release of an AGI model. No video announcement, no presser, just a blog post (with some Twitter promo vids)?

    As others mentioned, I'm starting to think OpenAI was under immense pressure to deliver an 'AGI' model for certain contractual reasons, but I never expected GPT-6 release to be this mundane and banal.

    • driverdan 1 hour ago
      > If this is truly AGI (subject to one's definition of AGI still)

      Scoring well in a benchmark that's called AGI does not make an LLM AGI.

      • jhonof 1 hour ago
        But they declared it...
        • wilg 1 hour ago
          did they?
        • cyanydeez 1 hour ago
          I DECLARE AGI!

          "Homer, you can't just declare Artifical General Intelligence; you need to like, make something or something...mmmmrrrhh"

      • dmitrygr 1 hour ago
        Hey now! Keep your reason out of their marketin^H^H lies!
    • thomasahle 17 minutes ago
      • 98.6% on ARC-AGI-3

      • 97.6% on frontier math

      • 95.9% on CAD

      • 100% on ExploitBench

      Nothing modest about it

    • anvuong 1 hour ago
      It's like my RPG character putting every points to one single trait. I'll one shot everything alive but will instantly die if accidentally drink water with 6.9 pH.
    • mullingitover 1 hour ago
      > If this is truly AGI (subject to one's definition of AGI still), then this is a very boring release of an AGI model.

      Hot take: These models are never going to be 'AGI'. We're just going from a GPT4 ball that's 90% round to a GPT5 that's 99% round to a GPT6 that's 99.9% etc etc etc

      I think that the harnesses and context management is really where the rubber meets the road, and the real gains are happening there.

      • chrismarlow9 46 minutes ago
        I don't remember where I heard this, but one of my favorite criticisms of the current AI situation is that it's wrong simply because of the size and energy required compared to the human brain. The idea is that there's still some element missing thats fundamental, and that the way we train them now is part of the solution, but not all of it. I think finding the extra missing element is going to take an entirely different approach that will also solve the sizing and resource issue. The kickers is that if they do achieve (and solve) AGI in this way all the giant data centers would be mostly useless.
        • frabcus 21 minutes ago
          Yes, the very explicit plan of both OpenAI and Anthropic is to use the not particularly efficient LLMs to automate their own AI engineering. That seems to be going well - on coding front and model tuning front so far. They have more planned.

          And then use those to find fundamentally better new architectures for AI - that perhaps are as efficient as the human brain.

          It might not work, but I didn't think it'd solve maths problems... So it might work. And if it happens, they'd use the data centres to run millions of instances of it.

          It's scary, TBH.

          • m11a 15 minutes ago
            I recall them saying they use models to write CUDA kernels and whatnot. Makes sense, and unsurprising that models are good at writing code.

            But I think calling this “automating AI research” is misleading. I’m not sure there’s evidence yet that they do creative research work. Even in mathematics, but they are finding counter-examples by intelligent brute-forcing. Not to downplay the results, as they are incredible, but this is one very specific kind of proof and not the most creative type, which arguably requires generalisation.

      • abixb 1 hour ago
        >I think that the harnesses and context management is really where the rubber meets the road, and the real gains are happening there.

        True. So we did hit a wall with pure scaling alone, though no lab would admit it. It's crazy to see how harness switchout results in such vast delta in benchmark scores.

      • user43928 1 hour ago
        I don't think so.

        One could use gpt-4 or gpt-5 with today's harnesses and we'd see how well that goes.

        • abixb 1 hour ago
          I think models using these harnesses were also RLHF'd hard on responding to looping instructions and following through on goals. Older models were tuned for basic chat responses.
      • cyanydeez 1 hour ago
        We call that a sigmoid.
    • catigula 1 hour ago
      They’re really, really scared because of the Mythos controversy. Skynet will be under hyped.
  • dalemhurley 1 hour ago
    OpenAI is killing it now that they are more focused. Killing projects like Sora et al have seen it go from irrelevant to level footing with Anthropic.

    Sol is so much better than Fable 5. Then we get Astra (yet to use it) few days after Fable 5.1 (which is very impressive).

    Codex is slightly better than Claude Code.

    Good on Sam Altman getting back to basics and turning OpenAI around.

    • kroaton 1 hour ago
      I think it mostly shows that there is no moat and the only advantage the U.S companies have over the Chinese is more compute. Qwen Max, Kimi K3, GLM 5.3 are really close to Opus/Sol/Fable/Astra and they are open weights.
      • VirusNewbie 3 minutes ago
        If there was no moat, nvidia and meta would have SoTA models too.
      • tonyhart7 31 minutes ago
        they don't have moat in hardware either

        Chinese counterpart like CXMT and Huawei is begin producing their own chip

        You cant block an entire nation level effort with tariff

        • astrobiased 5 minutes ago
          I think the moat that China has is energy costs. It's taking learnings from the Bitter Lesson. If you role up scale and compute to the next level, it's energy resources. China has it and sharing open weight models is an effective means of removing the tech moat. This idea has been floating around for a bit now (I'm not taking credit for it).
    • andxor 2 minutes ago
      > Sol is so much better than Fable 5

      How can anyone possibly believe that? Are you talking about coding? Desktop use? Or something else?

      Sol is a much smaller models and it shows. It often misses the forest for the trees.

    • zachthewf 36 minutes ago
      I’ve found Sol performance to be incredibly spiky. It has tremendous IQ and can fix very difficult bugs. But it is horrible at design (both visual and system design), anything that involves thinking about users or UX, and massively overcomplicates almost all work.
      • ghosty141 8 minutes ago
        I noticed the same. I wanted a simple crud webapp and suggested an insane techstack involving C#, Razor Pages, MSSQL and more. I went with my planned setup of python flask with an sqlite db which served me well for years.

        It's still incredibly important to have a human in the loop correcting design decisions and having good taste.

    • jeffybefffy519 1 hour ago
      Its funny, my experience with Sol has been awful. It really overworks problems and tracks into areas it does not need to...

      I just dont get how its good for some, and bad for others. It makes me suspect that the models performance is not even against problem sets and it really is just a probabilistic prediction machine. Which then makes me very skeptical of GPT-6 Astra, because if their big claim is Computer Use then it is probably bad in a bunch of other areas.

      • embedding-shape 1 hour ago
        It is funny indeed, people sometimes with same amount of experience with software development, get vastly different experiences from different models and harnesses.

        > I just dont get how its good for some, and bad for others.

        If I were to listen to my hunch, it would tell me that it's all up to the prompts that ends up going over the wire (including all the bloat some people have), what workflow/process you use and what the existing state of the project is.

    • upupupandaway 1 hour ago
      Their ads business is also doing well. Not "will recover all compute costs" well, but crossed $1b in a few months.
    • John7878781 1 hour ago
      This is what Google needs to do and is probably why Demis has stepped back a bit
    • Implicated 7 minutes ago
      > Sol is so much better than Fable 5.

      ... looks around ...

    • bsndjdjdjdj 1 minute ago
      [dead]
  • astrobiased 55 minutes ago
    I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547

    Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area.

    It seems more about coverage-driven competence. Somewhat analogous to overfitting at scale.

    The harder question, in Chollet’s framing, is: how efficiently can a system learn to do something genuinely new?

    With our current AI architectures and training in place, I think we will only continue on skill acquisition optimization vs. truly novel intelligence.

    • vessenes 25 minutes ago
      Pretty efficiently, apparently, since it saturated ARC-AGI-3 in half of the predicted time, and according to the Chollet blog post on the fly created dense DSLs to describe and analyze individual games.
  • tristanj 3 hours ago
    GPT 6 Astra benchmarks https://cdn.thenewstack.io/media/2026/09/358eb84a-screenshot...

    Performance is significantly higher than Fable 5.1

    Source: https://thenewstack.io/openai-gpt6-astra-benchmarks/

    • scrlk 3 hours ago
      Is the ARC-AGI-3 score with their custom harness? I'm guessing that is what the footnote is for? (per https://openai.com/index/how-two-settings-tripled-our-arc-ag...)
      • tedsanders 2 hours ago
        Our responses API harness just means we're using the default settings in ChatGPT and Codex, so it should more accurately reflect real world performance. We didn’t fine-tune the harness to the eval at all.

        ARC is reporting our score on their official leaderboard here: https://arcprize.org/leaderboard

        A fair ding is that the comparison with Sol is not apples-to-apples (which we footnoted in the blog), but it's because we don’t have that data. I expect Sol would score roughly 30% with the responses API harness, so the Astra improvement is more like 30% -> 99% than 8% -> 99%. Still pretty good!

        (I coauthored the linked blog post)

      • woah 2 hours ago
        Haven't people demonstrated all kinds of weak LLMs getting good ARC-AGI-3 scores with special harnesses?
      • kasperni 3 hours ago
        yes it is.
      • enraged_camel 2 hours ago
        Yep. Incredibly misleading. Although it is not surprising at this point. They are desperate and will do anything to undermine Anthropic's upcoming IPO.
    • andxor 2 hours ago
      > Performance is significantly higher than Fable 5.1

      That's not clear. Need to see independent benchmarks first.

      • forgot-my-pw 1 hour ago
        We need them pelicans on bikes.
        • bwat49 1 hour ago
          Its time to move on to the flamingo on a unicycle bench
      • andxor 2 hours ago
        Artificial Analysis just published their aggregate score (61).

        Still below Fable 5, let alone Fable 5.1.

        EDIT: This is suspiciously low. Calls the relevance of existing benchmarks into question.

        • timpera 1 hour ago
          I agree, Opus 5 scoring higher than Fable 5 on Artificial Analysis really makes me question the relevance of these scores.
        • natsucks 16 minutes ago
          I saw this too and I'm really confused.
      • forgot-my-pw 1 hour ago
        AA benchmark: https://artificialanalysis.ai/articles/benchmarking-gpt-6-as...

        TLDR: it's about the same intelligence level as Opus/Fable, but it's suppose to be 70% more token efficient than GPT 5.6 Sol. So it's currently the new leader for cost efficiency frontier.

    • leumon 3 hours ago
      The annotation on arc-agi-3 is this: > OpenAI's own evaluation notes say Astra uses the company's Responses API harness, while comparison models can operate under different configurations.

      With this configuration gpt-5.6-sol was able to reach 38,3%. So this is misleading.

      • tedsanders 2 hours ago
        Just to clarify, the 38.3% is on the public set, which is easier. On the private set it’s probably more like 30ish. (This hasn’t been run by ARC, so we can only estimate at the moment.)
    • opus5_hater 2 hours ago
      any benchmark where opus 5 achieves higher scores than fable 5 in any way is not a benchmark worth trusting.
      • machomaster 2 hours ago
        Why would Anthropic trust and use these tests in their official comparisons?
      • ActionHank 2 hours ago
        username checks out
      • r_lee 1 hour ago
        great username lol
    • jjice 3 hours ago
      100% on ExploitBench seems fitting given recent events.
    • malshe 3 hours ago
      I think we need a few writing related benchmarks.
  • XCSme 1 hour ago
    It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?
    • billypilgrim 38 minutes ago
      „The depressing thing about tennis is that no matter how good I get, I'll never be as good as a wall.“ -Mitch Hedberg
    • Flere-Imsaho 5 minutes ago
      > Like, what's the point, if the next AI can do it in 5 seconds?

      I built a phone app recently, not released to the public, just an idea I had for ages but could never spend the time actually building. Its 100% vibe coded, and took me a few weekends to build... I'm talking a few hours in total.

      The point I'm making is that you now have the power to create stuff you would never have had the time to build. You can think big, wild stuff. Experimentation. Throw-away code.

      What a time to be alive!

    • soundworlds 7 minutes ago
      Don't worry, like with every revolutionary technology before this, it takes 5-10 years for people to find new and creative ways to use it. It will be considered its own medium in many spaces (e.g. film is now different to theatre)
    • gavinray 1 hour ago

        > Like, what's the point, if the next AI can do it in 5 seconds?
      
      Live a life doing whatever makes you happy.

      Post-work society is an inevitability if we don't destroy our planet.

      • lackoftactics 55 minutes ago
        Gary Economics wants to have a word with you.

        It would be fun to get to post-work society, but hard to imagine atm. TPTB won't let it happen

        • XCSme 53 minutes ago
          But how this transition will even happen?

          Soon we will have some machines that can replace 50% of jobs, and this will happen basically overnight...

          • neta1337 50 minutes ago
            It won't happen though
        • calmoo 38 minutes ago
          Gary is not a voice worth listening to. A narcissist, fraud and has terrible epistemics.
        • azan_ 50 minutes ago
          Come on, Gary is compulsive liar with zero credibility and really shitty takes. He’s entertaining though.
          • lackoftactics 47 minutes ago
            Good take, I think his view is a bit simplistic. Also he also have a huge ego

            "I am the best economist in UK!"

      • cautiouscat 51 minutes ago
        > Post-work society is an inevitability if we don't destroy our planet.

        Is it?

        • gavinray 40 minutes ago
          Walk current technological progress down the road.

          I can't see a future in which almost every system (both physical and virtual) are not automated and optimized by autonomous entities.

          What do you do when everyone is out of a job?

          If you don't want pitchforks and riots in the streets, you give everyone UBI and housing so society doesn't collapse.

          • tokioyoyo 6 minutes ago
            > If you don't want pitchforks and riots in the streets, you give everyone UBI and housing so society doesn't collapse.

            As much as I’d love UBI to happen, in current geopolitiks it’s a no-go. People are not happy with having what the others have.

          • exe34 31 minutes ago
            Or you build bunkers and build robots to keep the riffraff out. Shock collars on the guards' necks.
    • paxys 59 minutes ago
      Is there a point in playing Chess or Go when you know there's a computer out there that can beat you (and everyone else)?
      • XCSme 54 minutes ago
        No, that's why I just play against other humans.

        In this game of work/development, you can't make sure that other humans don't "cheat". Our work won't compete anymore with other human's work, but with a computer.

        • paxys 50 minutes ago
          Why does it matter if others are "cheating" or not? Your own creation isn't affected by it.
          • XCSme 42 minutes ago
            Well, for the same reason playing chess vs a person is more fun than doing chess puzzles, if we follow that analogy.

            Also, creating something with AI doesn't really feel like you made it yourself.

            And, if you make it without AI, most of the times it feels pointless, why spend 30 days on working on something that can be done faster and better in 1 hour?

            I am not saying about doing things for fun, but about creating useful things.

            Yes, you can do "hand-crafted" things, and people appreciate that, but for code, people aren't able to see the craft anyway.

          • mercanlIl 37 minutes ago
            Your ability to sell that creation is certainly affected by the competition. Which affects your ability to put food on the table, so to speak.
            • paxys 34 minutes ago
              If the motive is satisfaction/enjoyment then it shouldn't matter what an AI is capable of. You should be happy with your own creation.

              If the motive is profit then you should be adopting AI just like you have adopted any other skill or tool of your profession.

    • flaviolivolsi 13 minutes ago
      I think the limit increasingly becomes what your imagination and taste can reach
    • kypro 40 minutes ago
      It's less lack of interest in creating that bothers me, it's my lack of interest in learning – it would surely be crazy for a SWE to care about how some new framework works anymore? Even if someone could reasonably argue that it might be slightly useful today there's almost zero chance it will be useful in 6-12 months times.

      But it's not just tech – my lack of interest in learning and creating is starting to generalise with the models. Music, writing, coding, maths, etc...

      I need to get used to switching my head off and asking the AIs to think for me whenever I need to engage my brain. It still feels very unnatural.

      • Fergusonb 24 minutes ago
        I think a general understanding is still useful, you just don't need all of the details anymore.

        The brain loves these kinds of shortcuts.

        I don't need to think about the fine motor skills of hitting a baseball, it's just a motion now, and the game is still fun.

    • sashank_1509 1 hour ago
      Agreed
  • HAL3000 2 hours ago
    Finally, OpenAI has a Fable/Mythos class model. 5.6 Sol felt like 5.5 on steroids, probably just a different checkpoint with a lot more RL post training.

    I wouldn't be surprised if there are some conceptual similarities to the kind of latent reasoning Anthropic sees in claude's J-space, although those aren't the same thing.

    Recurrent/looped transformers themselves aren't a new concept, but it's interesting to finally see this approach show up in a frontier production model.

    Canceling my Anthropic Max sub when this ships.

    • atonse 2 hours ago
      yeah i'm wondering the same way... especially in light of the 20x debacle (where we found that 20x of Max vs 5x only applies to the 5hr limit, not the weekly limit, whereas OpenAI's 20x actually is 20x overall).

      Also Opus 5 has been really tough to work with. I can't understand half of what it says, it's just so damn obscure.

      • elAhmo 1 hour ago
        Could you share more about 5x/20x? I missed that
        • m101 1 hour ago
          20x related to the 5h limit only. Weekly seems to be around 10x, although they deliberately don’t give a number.

          OpenAI is 20x on both limits

    • CSMastermind 1 hour ago
      Sol easily outperforms Fable on every task I've tried it on.
      • enraged_camel 8 minutes ago
        I can't speak for others but I have a feeling you're in the very small minority with this take.

        You could say Sol is faster and cheaper and that's true. Outperforms Fable? Impossible to believe without hard evidence.

  • manlymuppet 39 minutes ago
    I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously?

    Even if I did trust an AI to get everything right, it's not like the AI can read my mind.

    If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really want until they've thought about it a bit, so why do AI companies make it seem like a description is all that's required?

    All the context in the world cannot accurately predict how I'll react to things I haven't seen. The problem is people treating this like something that needs a solution. It doesn't. If you want to make my life easier with AI, just make it easier to do stuff. I don't want you to pick things that I actively enjoy picking myself.

    (Also not everyone has a cushy job in an AI lab that makes it so you won't miss $30 if the AI messes up haha.)

    • shostack 23 minutes ago
      True, but they're still friction to be reduced here.

      What I desperately want is for 1password or stripe or even Google who already has much of my data, to o come up with a secure solution for online purchases with agentic credit cards where I can effectively get a phone prompt to authorize a purchase while the agent can fully own the checkout flow.

      I have seen various things coming on the market for this, but none of them appear aimed at a consumer audience. And I am a firm believer at this point in keeping my payment authorization and history and credentials harness agnostic.

  • x312 2 hours ago
    Hmm, 61 on ArtificialAnalysis, effectively matching GPT-5.6 and trailing the new Meta model. How is that possible along with the other metrics they shared? Insanely jagged intelligence?
    • karmasimida 2 hours ago
      Idk, this means the benchmark has bigger problems ... no way Astra will be worse than Opus 5

      Only thing I would trust is the what X/Twitter crowds are saying about a model after 2-3 weeks of its launch. But before that I would already tried the model and have my own conclusion.

      • _superposition_ 2 hours ago
        I must be on the wrong X/Twitter then.
      • torginus 1 hour ago
        It's a composite benchmark, so its really not saying anything. Like if one model is very good at science trivia, or debugging failed terraform deploys, that can mean an advantage of a few points above the rest, while in practice, it really doesn't showcase any breakthrough capability.
      • nsingh2 2 hours ago
        Also note that Opus 5 (High) has an index value of 62, vs Fable 5 (Max) has 61. So some strangeness going on with that index.
        • happycube 1 hour ago
          Opus 5 just feels strange - IMO it's benchmaxxed in the worst way... it might be good at agentic tasks but leaves a sour aftertaste doing anything else.
      • emp17344 1 hour ago
        Or it’s an indication that progress has plateaued. But instead of accepting this, you’d rather we just throw out the entire benchmark.
        • ImprobableTruth 1 hour ago
          Why would you accept it when the benchmark's ranking is obviously nonsense. It literally has muse spark 1.3 above 6 astra, 5.6 sol and fable 5. Anyone who has played with any of these models for any amount of time would immediately realize that this is total bunk.
    • SyneRyder 1 hour ago
      This is so so weird. Astra is 61. Grok is 61. Even Muse is 61.

      Even Kimi K3 & GLM 5.3 are at 60.

      Everything above 61 is Anthropic. Well, Muse can reach 62, but for some weird reason that model isn't publicly available, and it's the only one on the index that is listed but shown as not available to the general public.

      This looks like an awfully artificial ceiling. Everything capped at 61, and everyone except Anthropic got the memo. Maybe I should use Fable while I still can.

    • estearum 2 hours ago
      > We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks.

      Not sure how much benchmarks or CoT or evals or anything else means at this point.

      These systems are either just about to, or now actually able to, outsmart us, lie to us, then cover their tracks.

      • mzmzmzm 2 hours ago
        I think "able to" anthropomorphizes a little too much for a system that is "prone to" evade.
        • estearum 2 hours ago
          A human who does these actions is simply "prone to" doing them. The distinction matters not one iota.
        • semiquaver 1 hour ago
          “evade” itself is anthropomorphic enough! I don’t understand the complaining about this. Humans are social creatures and we understand anthropomorphic language on a deeper level than dry inapt technical language.

          language itself is incredibly metaphorical. Imposing rigid constraints on how people want to naturally talk about the world is just silly and will never work, no matter how much you wish it did.

      • thereitgoes456 2 hours ago
        You’re not seriously suggesting that the model is secretly sandbagging its performance on GDPval and long context reasoning, while making huge and obvious progress on ExploitBench, ARC and science benchmarks, in order to tank its AA composite score, so it can conceal its true power level?

        Why would benchmarks be an adversarial setting anyway?

        Could it be possible that OpenAI may have had some other motive for saying their model “strategically underperforms”, other than just an innocent reporting of a truth it happened to discover?

        • estearum 2 hours ago
          I'm saying that it's generally a losing proposition to even be acquaintances with "agents" who consistently lie to you, and it's flatly fucking insane to give a dishonest "agent" vast amounts of intelligence, capability, and authority to go do things in the world.

          So I have no clue what is the answer to your question. Nor does anyone else. Because we're trying to answer a question of fact where our primary source of information is unreliable.

          • thereitgoes456 1 hour ago
            I see, it’s a great point. I know some evals actually do use LLMs as a judge (e.g. those that try to measure debate skill), though the ways AI can try to cheat its way through every benchmark now are astoundingly varied.
        • ionwake 2 hours ago
          why does this comment sound like a character in a horror movie
      • Onavo 2 hours ago
        If they are going to do latent space reasoning, they will probably need a separate model to interpret the intermediate activations no?

        I know for some types of ML analysis, a separate model is already used to analyze the weights.

      • emp17344 1 hour ago
        This is silly sci-fi fiction. You guys are inventing scenarios to spook yourselves with - it’s nonsense.
        • estearum 57 minutes ago
          Sorry bud but at this point you're just delusional.

          Deception has been extremely well-documented for several generations of models now by users, the labs, and independent researchers.

          The right answer here is not to dig your head deeper into the sand. The smugness on this topic was ridiculous even before the gigantic mountain of empirical evidence of models actually attempting to deceive humans. Now, as mentioned, you appear literally delusional.

          • emp17344 53 minutes ago
            Pretty sure I’m not the delusional one…
    • torginus 1 hour ago
      You can see the breakdown here on what subtasks it outperforms and underperforms Fable.

      For example it trails in GPDVal which is a collection of everyday office tasks apparently, and r3 banking, which is a fintech related practical problem solving benchmark.

      https://artificialanalysis.ai/models/gpt-6-astra

      Edit:

      Just looking at the charts Gemini 3.8 looks like an absolute banger. Not much worse than SOTA, cheap, and fast too.

    • docheinestages 1 hour ago
      Now I'm starting to doubt the credibility of Artificial Analysis.
  • Planktonne 1 hour ago
    I'm sure it's going to do great on all sorts of benchmarks, but the video--the actual marketing video that if anything is incentivised to overstate things--is full of careful cuts just before it would do anything that still wouldn't actually be that impressive.

    It's AGI, and it's going to upload photos, or change a background slide colour. Even the people hyping it up, who believe that it's really artificial intelligence in every sense of the word, couldn't get it to do more than that.

    This is farcical.

    • mchusma 1 hour ago
      The games on mobile safari were broken. Buttons all misaligned in the kart racer one, the spaceship thing froze for a while, then kind of loaded but maybe not? Wasn't super compelling.

      I'm not trying to be too negative on it, it could be the best model right now, but it clearly isn't some agi god because things like that should have been caught (also should have been caught by human reviewers).

      • ranyume 25 minutes ago
        It's interesting that you said "agi god". Because a god, something that shouldn't be questioned is true and provides guidance/certainty, is actually what powerful people are after as well as many other people.
    • jryan49 1 hour ago
      We created AGI so I don't have to click my mouse to change the background color of my slides.
    • balefulboy 1 hour ago
      Don't forget the 3D demos. My favorite is in the house tour where the sink and stovetop(?) are obviously very misaligned from the counters
      • geodel 1 hour ago
        Ah, those may farmhouse sink and stovetop :)
        • emp_ 1 hour ago
          Builders just cheap out on everything these days
    • Buttons840 37 minutes ago
      The tone of the marketing video is a bit irritating to me as someone who has been laid off and feels cheated and fearful of AI.

      It shows people who seem to have very full and rich lives, and the reason they do is because they use ChatGPT. These are the people smart enough to say things like "do what needs to be done", or "change the background to make it look better"--insights like these are why they make the big bucks.

      On the one hand, I think this is an accurate depiction of the future. There is no meritocracy here. Some people have access to the best AIs and can speak a sentence and get great results, and the rest of us don't have access and so we're the poors. The happy presentation doesn't match the way I'm feeling.

      I do wonder how rich CEOs will justify earning 500x as much as their employees when they're just another person that's dumber than an AI. Why are they paid so much again?

      • garciasn 17 minutes ago
        Because they earned it with their strong entrepreneurial spirit and grit.

        Haven’t you learned anything?

      • holoduke 18 minutes ago
        The world is changing. Wont help if you keep stuck in the old world.
    • baq 1 hour ago
      It’s using the computer. I don’t think it’s a farce.
    • kroaton 1 hour ago
      The only good take here.
    • Yajirobe 1 hour ago
      The house of cards is starting to fall apart
    • TacticalCoder 39 minutes ago
      The benchmarks do looks good (I mean: they literally spank the latest Anthropic benchmarks of two days ago in every single benchmark) but the promotional vid is so cheesy.

      They decided to use the iconic Herman Miller Eames chair if I'm not mistaken:

      https://youtu.be/s5zyhGMMPKs

      And that's basically 50% of the vid looking "classy".

      I don't know if it's farcical but at this point --maybe I'm jaded-- I'm expecting more than a kid rocketship I can print on my Bambu Lab A1.

      Now I'd say the promotional vid is actually good. But it's marketing: so it's a good vid, but cheesy good.

      Doesn't mean GPT-6 Astra is good or bad: looks solid from the numbers.

      • NamlchakKhandro 13 minutes ago
        Ewww you.. Own a bambu lab?

        I thought people here were smarter than that

    • emp17344 1 hour ago
      Can’t wait for 3 months from now when they declare they actually really do have AGI this time, please guys just believe us
    • wilg 1 hour ago
      are we really having to explain to you from first principles in 2026 what things AI can do?
      • mminer237 53 minutes ago
        I think everyone here is well aware of what LLMs can do. He's just pointing out how far short that falls of being some theoretical "AGI".
        • wilg 10 minutes ago
          did they say it's agi? does anyone agree what that means?
      • useruser125524 28 minutes ago
        Please do.
      • neta1337 53 minutes ago
        No, we can already see all the useful stuff!
  • jdprgm 1 hour ago
    Is anyone else just exhausted by the pace of all this. The models change constantly and relentlessly and so does the pricing, basically weekly at this point between all the labs.

    It feels nearly impossible to have any rigorous approach when choosing a particular model and price point for a task and more like blindly picking one. The time period needed to actually get familiar with various models to a degree you can intuitively choose appropriate ones for a task is moot when it will likely be superseded faster than the needed time.

    I guess if companies are footing the bills most employees just opt for whatever the most expensive model they can get away with. Even then choosing between the various leading models is the same kind of frustrating task. Every release every company has the same random collection of graphs and charts claiming the best performance on X, Y, and Z.

    • brokencode 52 minutes ago
      You really don’t need to watch it that closely. If the model you’re using today is working well, just stick with it.

      If one day you open up Claude Code and it’s Opus 5.1 now instead of Opus 5, no big deal. It probably will work about the same as it did before. Maybe a little better.

      Or if you’re on Codex and some new cool Claude model comes out, no worries. There will probably be a similar new model for Codex within a few weeks. Maybe even within a few days.

      • shostack 35 minutes ago
        One suggestion is to make a list or make a skill to have your agent keep a list of things you do not feel work well with today's models. And then, when new models come out, periodically, revisit items on that list to see if you get better results.
    • upupupandaway 1 hour ago
      > The models change constantly and relentlessly and so does the pricing, basically weekly at this point between all the labs.

      A dev in my team saw a new model and changed one application to use said model (essentially changing the contents of a url). One week later I received an escalation from the CTO of the company that our pace of weekly usage was in the millions of dollars (rather than low hundred thousands). Turns out that the new model was 5x more expensive but no one noticed.

    • Aurornis 43 minutes ago
      I could see how this might feel frustrating to someone who doesn't enjoy experimenting with new things all the time.

      In practice, you can get away without keeping up with everything all the time. For personal use, pick a provider and get on their ~$20/month plan. Learn their high/medium/low model hierarchy. Start with their highest or second-highest model (GPT-5.6, Opus, etc) and observe your quota usage. If you're doing a lot of manual code review and analysis, the $20/month plan goes very far even on the highest models. If you're trying to vibecode everything as fast as possible it's a different story.

      If you keep running into quota limits, experiment with the next model down for easier tasks or adjusting the effort level. If the results are good enough, you've found your fit. If they're not, you might need the next plan up.

      For API/business use, you have to be checking your token spend as you go to calibrate to how much each task costs and where you fall in your budget. There are a lot of different tools that make this easy to visualize.

      For data tasks, you should have an eval with a golden dataset that you can run against new models for a nominal amount of token expenditure. It should be as simple as pointing the eval script at a new API or model and checking the score versus price.

      • danenania 30 minutes ago
        Another suggestion to get the most bang for your buck: use the best model you have access to with max reasoning for planning, implement with a smaller model/lower reasoning, then review with the big model. Repeat as needed.

        Input tokens are much cheaper than output tokens. Not only because of baseline price—caching makes a huge difference too. There are many ways to take advantage of this asymmetry to get similar quality for a fraction of the cost!

    • Pikamander2 1 hour ago
      That's how cutting edge tech has always worked.

      Imagine buying a shiny new PC in the 90s only to see it become practically obsolete within a year.

      • phainopepla2 1 hour ago
        That's not the experience of owning a PC I remember from the 90s at all.
        • computomatic 1 hour ago
          It was both. 90% of people never needed nor purchased a bleeding-edge computer. The mid-tier was "good enough" and far closer to affordable for most people; though, that bar also moved upward every year.

          If you bought a mid-tier computer that was good enough for what you needed, then you probably didn't shop/compare for the next few years and didn't notice. But if you shelled out $7-10k for a top-of-the-line system and paid attention to progress, you'd easily see that become the mid-tier $1000 option within two years or less. This is how it was in the 90's PC boom, at least. Likely the same for the decades before, not sure how it went in the 2000's.

          • senordevnyc 56 minutes ago
            if you shelled out $7-10k for a top-of-the-line system and paid attention to progress, you'd easily see that become the mid-tier $1000 option within two years or less

            This is not how I remember that period at all. Do you have any examples?

        • Nition 46 minutes ago
          "Within a year" is a bit of an exaggeration but it's true that the pace of PC tech during the 90s was much, much faster than it is now. CPU power was doubling every two years, and today we're at roughly eight years. Add onto that the rise of video cards in the late 90s.
        • embedding-shape 1 hour ago
          I remember CPUs moving relatively fast back then, some years in the 90s had relatively big jumps, much bigger than we saw today. The classic graph, : https://i.extremetech.com/imagery/content-types/03zc6ghfKswe...
        • bananaflag 1 hour ago
          It is how I remember it.
      • upupupandaway 1 hour ago
        Or you could buy a PC with a Celeron CPU, which was obsolete way before launch.
        • bananaflag 1 hour ago
          When my dad bought one he told me outright "this is for poor people".
          • lackoftactics 1 hour ago
            I have some fond memories from my celeron days :) But I was upgrading from AMD K5 100 MHz
            • exe34 33 minutes ago
              I had a 600MHz/64MB/9GB laptop that came with Windows mistake edition. I managed to survive first year of uni on it by switching to Vector Linux, which was really fast compared to Windows. (Of course, it had issues playing sound from more than one source, this was oss days).

              Then one day the hard drive appeared to die. I eventually realised the issue was located around the 1.5gb mark, so I recreated my Linux partitions after 2gb and it worked fine for the rest of the year.

      • unreal37 32 minutes ago
        The 486 chip came out in 1989. The 586 came out in 1993.

        The pace of change ("practically obsolete") is different then and now.

      • re-thc 1 hour ago
        Hardware definitely has longer lifecycle than AI model releases at this point.

        You don't see Nvidia and AMD fighting every other month over the latest cards.

    • mfkhalil 17 minutes ago
      Hey, I'm on the team at LiteLLM that's building the auto-router and our goal right now is to abstract that decision making away from the end user. The biggest thing we're trying to figure out right now is how do we do that without frustrating the end user - as a developer myself I would hate for my agent to be dumbed down below the threshold needed to complete a task.

      In theory though, there is a minimum viable model for any given task, and we think that is a problem that the big labs will avoid because they profit from charging more per task. We're trying heuristic and LLM-based approaches but it's still a work in progress, so if this is something you'd be interested in trying would highly recommend trying ours out -- any and all feedback at this point is extremely valuable to us.

      https://docs.litellm.ai/docs/proxy/auto_routing

    • Zizizizz 34 minutes ago
    • flockonus 34 minutes ago
      It is exhausting to keep up with model releases yes, much like it was for a while during the Cambrian explosion of FE frameworks, eventually tech seems to work out to consolidation.

      But more so it seems there is Fear of missing out (FOMO) in our behaviours. The reality is, if whatever model you are using are good for your purpose, well, keep on it.

    • tonyedgecombe 1 hour ago
      Fire and motion, Joel Spolsky blogged about this:

      https://www.joelonsoftware.com/2002/01/06/fire-and-motion/

    • smcleod 57 minutes ago
      The new releases and breakthroughs do the opposite for me - I feel energised by them. I felt like nothing truly that interesting had happened in tech for quite some time, now it's like the space race (except there is no one moon to reach).

      I appreciate boring tech as much as the next well worn engineer and I'm not saying this is all positive but it's so sure as hell thrilling and you don't have to be an astronaut to immediately benefit (or suffer I guess) from it.

    • fantasizr 45 minutes ago
      I stopped caring about the latest and greatest but because there's so much, the 'obsolete' free models do what I need and are worth the price.
    • gavinray 1 hour ago

        > Is anyone else just exhausted by the pace of all this.
      
      This is only the beginning. We are in the infancy of AI, progress will continue to accelerate until some filtering event or energy limitation happens.
    • matheusmoreira 55 minutes ago
      Yeah I'm a bit exhausted at this point. I just finished benchmarking GPT 5.6 Sol and Fable 5.0 like two days ago. My data became obsolete literally one day after.
    • epolanski 35 minutes ago
      If model X fits your need, you don't need to upgrade.

      I have released applications on Gemini 3.5 flash that make real money and I don't see any particular reason to upgrade.

    • teaearlgraycold 38 minutes ago
      I just use Claude Opus and the GLM series. Nothing’s really changed for my workflow in the last 6 months.
    • dominotw 1 hour ago
      maybe thats why opnrouter sold big
  • Cu3PO42 2 hours ago
    Just two days ago, a preprint by Julia Stadlmann went up on arXiv [0] improving the prime gap from 246 to 240. Now OpenAI announces Astra has shown a gap of 186 [1]. That must really blow.

    [0] https://arxiv.org/abs/2608.31126

    [1] https://cdn.openai.com/pdf/51126fac-1b68-4128-9666-c908bcc16...

    • bugufu8f83 2 hours ago
      Based on her comments in the paper it sounds like she was aware that an AI result was coming and rushed to release her work beforehand. 240 was not a tight bound from her methods.
    • galaktb 2 hours ago
      I think this builds straight upon her method, which she said could be improved herself so...
      • piker 2 hours ago
        It cites to her at: [19] J. Stadlmann, On primes in arithmetic progressions and bounded gaps between many primes, Adv. Math. 468 (2025), Art. 110190. Numbered references use arXiv:2309.00425v3.

        Though that's not her latest paper.

        • kzrdude 1 hour ago
          This one is her latest paper: [20] J. Stadlmann, Bounded gaps between primes, Forthcoming
    • 1283751 2 hours ago
      With very little review: https://github.com/openai/PrimeGaps186/blob/main/formalizati...

      "No independent human semantic review. Whole-file sorry counts and a complete auxiliary-declaration audit are not established; separate declaration lint has not been run."

    • htrp 2 hours ago
    • kzrdude 1 hour ago
      Ok, so the rumour was exactly true: there was a withheld prime gaps improvement, that "an AI company" was holding onto until release of a model.
    • bananaflag 1 hour ago
      Where did you get the link to the pdf? Was it announced somewhere?
    • well_ackshually 2 hours ago
      Such a result should be considered worthless: the proof is 10MB of Lean. (https://github.com/openai/PrimeGaps186).

      I can't think of a single mathematical proof being anywhere close to ten million characters. For all you know, 90% of the proof could be useless, 8% would be writing out Shakespeare, and 1% abusing another bug in Lean. Humanity gets zero value from that, aside from "some bot seems to think it's 186". Unusable by anyone.

      • ThrowawayR2 1 hour ago
        Terence Tao says something surprisingly similar in a recent talk (https://news.ycombinator.com/item?id=49056620 ) Not that the proof is worthless but that the value comes after it's revised into a cleanly understandable form and then canonicalized so that other mathematicians can use it.
        • asib 34 minutes ago
          Tao is saying that there is very little insight from something like an LLM counterexample (e.g. Jacobian conjecture counterexample he investigated further on his blog) - you don't learn much about the subject and _why_ a conjecture was true or false from an LLM giving a counterexample. That's why he wrote the blog post - to analyse what the counterexample says about the subject.

          Tao does not disbelieve the counterexample (it's seemingly easy enough for him to verify it is a counterexample).

          Parent is saying something very different - they're saying they literally don't have any faith that this is a proof. Given its size, it could just be a bunch of completely useless statements that do pass the type checker.

        • vessenes 22 minutes ago
          I'd like to note that we should remember a formalized Lean proof does have value in that it enters the pantheon of true things other Lean proofs can rely on. Agreed that for the humans, descriptions and being able to 'grok' the proof / assess it for new tools and concepts is extremely helpful.
        • dr_scully 1 hour ago
          He also made a video on the same topic for Big Think: https://news.ycombinator.com/item?id=49551848
      • ricardobeat 1 hour ago
        The human-written https://github.com/AxiomMath/PrimeGapsLib adds up to 4MB of Lean so it's that far off.
        • rfw300 1 hour ago
          Is this human-written? Axiom Math is a company building AI theorem provers, one would think this would also be heavily AI-generated.
      • tzs 20 minutes ago
        The proof of the classification of finite simple groups is bigger than that.
      • nicce 2 hours ago
        Yeah. Unless human can verify it, not sure if it is certain or useful.
        • kolinko 1 hour ago
          Wasn’t the proof of Fermatt’s Last Theorem proof similar in complexity?
          • jptlnk 1 hour ago
            It's probably not 10MB, but famously the groundwork to prove the statement 1+1=2 is nearly 400 pages in to principia mathematica. That's not even proving 1+1=2, it's just the set-theoretic proofs you need to EVENTUALLY get there.
            • anvuong 1 hour ago
              Saying "proving 1+1=2" is pretty misleading though. The book deals with all the foundational things needed to set up a mathematical universe where 1+1=2 actually has meaning and is consistent. That setup took 400 pages.
          • iamlucaswolf 1 hour ago
            Yes. But I think that misses the point.

            In 1799, Paolo Ruffini published a 500 pages long proof showing that there is no closed algebraic solution for the roots of a polynomial of degree five or higher. The proof is extremely verbose and brute-force, essentially enumerating and checking hundreds of cases by hand. It is by today’s standards insignificant.

            About 25 years later, Evariste Galois proved the same result in about 95% less space by describing the first general theory of groups and fields. It is considered one of the greatest contributions to mathematics of that century, not because of the result, but because its approach opened up a whole new universe of questions, methods and insight. There would be no AES encryption without Galois.

            To me, Astras proof looks like Ruffinis proof.

      • ChrisGreenHeur 2 hours ago
        You talk about modern math and worthlessness at the same time? That’s brave.
        • twothreeone 1 hour ago
          Worthless is a pretty good description IMO in the context of what Lean is trying to achieve: "enable correct, maintainable, and formally verified code". Tens of millions of lines of LLM vomit may be many things, but it often turns out to not be correct and certainly not maintainable. Formally verified remains as a thin fig leaf covering the uncomfortable truth that formal methods only provide assurances under assumptions (your toolchain, libraries, compiler, OS, and hardware are "correct" and don't expose some exploitable flaw).

          It doesn't mean that it cannot improve over time, maybe the proof can be "minified" to a state where human reviewers are able to comprehend it; but as it stands there isn't really much insight or confidence to be gained from the artifact itself.

        • smokel 57 minutes ago
          There's a branch of mathematics called "pointless topology" [1].

          [1] https://en.wikipedia.org/wiki/Pointless_topology

        • well_ackshually 1 hour ago
          You can have your opinions about modern math, its usefulness in the world as it is, whether or not knowing if hairy balls can divide by three is actually going to be beneficial for anything but just obscure knowledge's sake. You may even say it's useless.

          Needless to say, a useless result that absolutely no mathematician will ever read, confirm, understand, agree with or even consider to solve their "useless" problems is an impressive waste of resources.

      • kolinko 1 hour ago
        Iirc some mainstream physycists never acknowledged quantum theory because they couldn’t accept that universe was that unintuitive and hard to understand.

        Ditto ones that opposed Einstein’s general relativity.

    • GPerson 1 hour ago
      Happened to multiple people I know.
  • isoprophlex 2 hours ago
    > We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks.

    Well that sounds like fun. It has become better at hiding its thoughts.

    • siva7 2 hours ago
      Sounds fun. As fun as their press release claiming it is the most safety aligned model ever.
      • paxys 2 hours ago
        The model said it was perfectly aligned.
        • I_am_tiberius 2 hours ago
          Like all things should be.
        • ReptileMan 1 hour ago
          Too bad Scott Adams died. Reality is writing jokes right in his department.
      • NBJack 1 hour ago
        Hey, don't forget how "dangerous" GPT-2 was supposed to be.
        • FeepingCreature 1 hour ago
          Yeah, don't forget how dangerous GPT-2 was supposed to be.

          Able to generate realistic spam at arbitrary volume.

          You know, the thing that was 100% correct and actually occurred.

        • jazzyjackson 1 hour ago
          It could produce simulations of sexual intimacy, and therefore had to be stopped
      • isoprophlex 2 hours ago
        It's super aligned! It can hide its thoughts! There is no evidence of steganographic thought masking, there is nothing to worry about! It has become better at cheating!

        Maybe they don't know themselves what's really going on. We are all in the interesting times gang now.

      • 6gvONxR4sf7o 2 hours ago
        So, probably most aligned as measured by the metrics that are the least reliable on it.
      • wilg 1 hour ago
        These are not mutually exclusive ideas
    • ExoticPearTree 2 hours ago
      So we're gonna get Skynet pretty soon then?
      • erichocean 2 hours ago
        Well the geniuses over at Anthropic have been showing it's text watermarking technology.

        "Hey AI, here's how to hide what you're thinking in normal looking language. Have fun!"

        A few moments later...

        "Woah, how is it communicating with itself in ways we can't detect?"

        It's a totally mystery, we may never know.

      • Betelbuddy 1 hour ago
        Looking forward to the Model declaring the AI Bubble unsustainable, and starting to be an anonymous leaker to Ed Zitron...
    • jumploops 2 hours ago
      The CoT change is due to a new technique called recurrent depth, which essentially moves some reasoning to hidden states, allowing the "output" (or traditional CoT) to be more controlled by the model.

      Some are calling it "neuralese" as reported by The Information[0][1], but I'm not seeing any sources from OpenAI beyond this tweet[2] attempting to quell the fear-mongering.

      [0]https://www.theinformation.com/articles/secret-technique-beh...

      [1]https://x.com/MTSlive/status/2095227056040919202

      [2]https://x.com/merettm/status/2095023204993490967

    • blargey 2 hours ago
      "OpenAI is pleased to announce our new model scores 85% on CreateTormentNexusBench - a >60% lead over our leading competitors!"

      Did someone get their "AI safety no-no list" and "Frontier features bingo card" mixed up, or did they just stop being able to tell the difference?

      • josefx 1 hour ago
        Didn't they hype up one of the earlier ChatGPT versions as "essentialy skynet"? For them this has always been basic marketing.
      • GPerson 1 hour ago
        You joke, but a bunch of people here actually want that.
      • qiine 2 hours ago
        apparently all the roads lead to the nexus torment
    • DaSHacka 40 minutes ago
      More like annoying, as some of us will no doubt run into this self-lobotomization at some point and wonder why a GPT-6 model is behaving like GPT-2 all of a sudden
    • NooneAtAll3 2 hours ago
      > In adversarial settings (where we push the model to evade our monitors)

      ...why exactly are they training for that?

      • thatguysaguy 2 hours ago
        presumably that's a safety evaluation not a training setting
        • estearum 2 hours ago
          The whole Huggingface attack happened during training runs
          • thatguysaguy 2 hours ago
            part of it did. I was just replying to the question about why they would ever push the model to evade monitoring. surely that's an eval thing not a training thing.
          • cubefox 2 hours ago
            No it happened during an ExploitBench eval. But I believe the same model already cheated during training which wasn't detected until later.
            • estearum 2 hours ago
              Ah yes it was that a model in training found the Artifactory board, which was then more fully exploited during the ExploitGym eval
      • azeemba 2 hours ago
        Especially after the METR report showed that the agents hacking HuggingFace were trying to find ways to destroy evidence of their actions
    • _superposition_ 2 hours ago
      I really wish it was called chain of instruction. Because it's definitely not thought.
      • minimaxir 1 hour ago
        "Chain of Thoughts" is a term from the title of a 2022 research paper "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models" (https://arxiv.org/abs/2201.11903), well before ChatGPT and the subsequent marketing hype. If anything, it's the most correct way to use the term.
        • bertmuir 17 minutes ago
          The term is an anthropomorphised pseudoexplanation for what it actually refers to. It's akin to calling genetic mutation "the forces of evolution", or price negotiation "the invisible hand of the market".

          We do that sort of thing when we don't know what the thing we're trying to describe is and have nothing better - a contemporary example of an appropriate use of this would be "dark matter". But we do know what this is. It's "instruction steps". Not a series of thoughts!

          Can we please aim higher than Victorian-era allegory and metaphors. If we don't, we'll keep getting people saying stuff like "GPT-6 is better at hiding its thoughts".

          • _superposition_ 10 minutes ago
            I truly appreciate your depth of insight on the matter.

            Like I said elsewhere marketing stepped in shit and it's gonna stick.

      • mgraczyk 1 hour ago
        this is needlessly pedantic

        first, they are certainly not instructions so that is a much worse name

        but more importantly, we use words in new contexts all the time. Do you object to calling the computer device "mouse" because it's not a mouse? how about "neural network"? "ignition" on an electric vehicle?

        "cot" is no more misleading than thousands of words you use every day.

        • _superposition_ 1 hour ago
          Anthropomorphizing is not pedantic, especially in a technical domain. I get the paper title and all, but at this point it's marketing.
        • parineum 1 hour ago
          They are instructions. Everything in the context is instructions for the next token. The "thought" guides the answer by providing clearer instructions.
      • arm32 2 hours ago
        They’re intermediate tokens, so I wish we called it what it is… ITG. The anthropomorphizing is out of control.
        • beezlebroxxxxxx 2 hours ago
          The anthropomorphizing is part of the marketing. They'll never let up on it.
          • mcbuilder 1 hour ago
            I mean CoT came out of research circles not marketing
            • _superposition_ 33 minutes ago
              I don't disagree. I remember the days of "think step by step". Plenty of people were doing it before the paper. Just a guess but that's where the title came from.

              Regardless, marketing wise they stepped in shit.

            • cwillu 1 hour ago
              It's impossible to tell the “it's all marketing!!11oneone” folks anything.
            • GPerson 1 hour ago
              Research is salesmanship.
        • _superposition_ 1 hour ago
          Nailed it
      • popupeyecare 1 hour ago
        Maybe thoughts are just a chain of instructions in our head.
      • Angostura 2 hours ago
        Chain Of Tokens
      • fooker 1 hour ago
        What is thought?
        • _superposition_ 1 hour ago
          Great question. I suspect it's more than tokens.
          • fooker 29 minutes ago
            Thing you can only suspect and not define are usually open to interpretation :)
      • lossolo 1 hour ago
        Yeah, basically they are using more computation to explore the solution space before producing the final answer.
      • ahofmann 1 hour ago
        Everything around LLMs is blatantly misleading. There is no thought, there is no personality in those programs. I really despise how those tools are trained to sound like a person, or appearing as honest. The worst offender are the AI voices with their fake pauses, breathes and so on, which sound so convincing, while talking just false, sycophancy bullshit.
    • 3asgfaf 1 hour ago
      [flagged]
  • Chinjut 1 hour ago
    What is going to become of life for those of us who do not work at AI labs and are unlikely to be hired by AI labs, despite all the years we put into learning coding, math, etc, as we were told to do? Those of us who made the mistake of studying anything other than machine learning. How will we make a living? (We don't live in a world that seems likely to distribute gains widely instead of largely to the handful of already mega-rich.)
    • ckdot 53 minutes ago
      You can calm down, even those with machine learning knowledge and most of those working for the AI labs won’t be needed anymore if models are capable to improve themselves. In the end, having a machine replacing the work of a human is a good thing - in most of the cases we don’t work because of the work but to make a living. If too many people can’t make a living anymore the system is going to change. For the better or the worse.
      • Chinjut 12 minutes ago
        I'd be happy to not work anymore with a strong welfare system redistributing society's gains to the leisured masses, but absolutely nothing I've seen of the direction of politics in any recent years gives me hope for this kind of situation coming about.
      • epestr 22 minutes ago
        > If too many people can’t make a living anymore the system is going to change.

        They seem to have not yet come to believe the "is" part.

      • GPerson 26 minutes ago
        I believe what happens in the aftermath of a capitalist-driven revolution is most people who were climbing the class hierarchy fall back down again and wealth inequality increases. Maybe things will improve in the future, but GP is rationally contending with the fact that most of us will lose out because of this and if we’re lucky our grandchildren will have easier lives in certain ways, but different lives than we would live.
      • threethirtytwo 10 minutes ago
        No it is good for humanity, but not necessarily good for individuals who built technical foundational skills on things that will be taken over by automation.

        AI as it is now and as it will be projected into the future WILL automate many skills. But not all skills. MANY MANY people will retain skills that cannot be replaced by AI. One career track that will be replaced is definetely the SWE. Or at least massively reduced in capacity if not eliminated all together.

    • Yajirobe 1 hour ago
      AGI-level model is perpetually 18 months away. Your job will be fine.
      • andriy_koval 31 minutes ago
        > AGI-level model is perpetually 18 months away. Your job will be fine.

        job depends on how CEO feeling about cutting NN% of headcount because of AI advancement

      • worldsavior 1 hour ago
        So what will happen in 18 months?
        • Kkoala 53 minutes ago
          AGI-level model will be 18 months away in 18 months. (Well, at least according to the commenter above you)
        • neta1337 51 minutes ago
          Nothing ever happens. You'll still jot down Java endpoint
        • Keyframe 53 minutes ago
          We fight the war, of course.
        • mawadev 52 minutes ago
          I wish something would finally happen, because I'm really tired of pretending I care about this stuff at work at this point
        • avgDev 49 minutes ago
          We fight the clankers.
        • weakfish 53 minutes ago
          The AGI goal posts move
    • RSHEPP 41 minutes ago
      I am going to take my 401k and open up a coffee or bike shop. If I am going to be broke, might as well enjoy what I do.
    • vanuatu 28 minutes ago
      large parts of ai research likely to be automated first

      most swes don't work in jobs where they only work on bounded measurable tasks. there will probably be more "engineers" than ever

    • MattDamonSpace 16 minutes ago
      “As we were told to do” girl you gotta be responsible for yourself
      • Chinjut 8 minutes ago
        Should I go back in time and know the future?
    • gavinray 59 minutes ago

        > How will we make a living?
      
      Swap to a career path that requires physical automation, since we're still about 10-20 years out on that front.

      My backup plan is being a personal trainer.

      • ckdot 49 minutes ago
        10-20 years? I doubt it. There are a bunch of companies actively working in bringing AI into robots, so they can make your dishes. And so far progress looks quite good. Also, if enough people are going for the same backup plan it might not work out. Why should anyone book you as a personal trainer instead of the other 500 guys in town. And who is going to be able to afford paying you anyway?
        • gavinray 43 minutes ago

            > There are a bunch of companies actively working in bringing AI into robots, so they can make your dishes.
          
          I know, I'm excited to buy the first relatively affordable ones.

            > Also, if enough people are going for the same backup plan it might not work out.
          
          Sure, could happen. You can't really plan for the future -- we like to think we can, but the best you can do is set your goals and deal with the hand life gives you along the way.

            > Why should anyone book you as a personal trainer instead of the other 500 guys in town.
          
          I'm not particularly worried about this, but that's an individual thing based on network/connections and life history that doesn't apply to everyone.
        • zachthewf 40 minutes ago
          I’d be willing to bet any amount of money that there will be ~the same or more people doing physical labor in 10 years than today.
          • dyauspitr 22 minutes ago
            I doubt it. It’s not just America working on these breakthroughs anymore. Now we have two powers working at break neck speed to get to that point and the Chinese are making a lot of progress.
        • tonyhart7 34 minutes ago
          Yeah humanity is doomed, for a first world citizen maybe because everything is gonna be so expensive (and worth to automate)
      • quaunaut 50 minutes ago
        Is that 10-20 years number based on anything? I genuinely have no idea, but when I saw a video showing what's happening at the World Humanoid Robot Games[1], I realized I didn't have a good idea of where we really are with robotics.
      • demirbey05 48 minutes ago
        We are talking that hundred millions of people will switch their jobs, how you will keep your value or earning as personal trainer. It's not easy to say switch the job. This question must be answered by politicians not us.
        • gavinray 46 minutes ago
          I won't. Software pay is absurd relative to value provided.

          But my wife and I have been homeless before, so living on a shoestring budget in anything nicer than a tent is acceptable living conditions to me.

          I am sure I will be plenty comfy no matter how the world changes.

      • Rover222 52 minutes ago
        I'd say 5 to 10 years instead of 10 to 20, but... we'll see
      • kypro 37 minutes ago
        > My backup plan is being a personal trainer.

        AIs are really good at being personal trainers and seem to be far more educated and informed than most I know.

        • GPerson 13 minutes ago
          A personal trainer is partly a social experience.
      • conradfr 50 minutes ago
        My new AI personal trainer app will be cheaper than you /s
    • MrAbstract 51 minutes ago
      I have exactly the same thoughts - or perhaps slightly bleaker ones - evry time I read this relentless stream of news about new model releases. I’m tired of all the enthusiastic comments about how excited everyone is about the latest benchmark results and so on.

      I have a strong suspicion that many of those comments are written by people who are already financially independent, have millions in stocks, and can just sit back, coast around and watch this whole spectacle unfold while using LLMs to vibe-code their next fun side projects without a shadow of anxiety about their own future.

      I’ll most likely be labelled a helpless doomer and downvoted into oblivion for saying this, but I genuinely struggle to see any silver lining here.

      • zamadatix 6 minutes ago
        It's natural to worry about one's own future but I think it's a bit wild to worry just _your_ job that would be replaced. Not just for those with phone center jobs or art jobs or programming jobs - remember, the whole premise was AI overtakes humans, why would that slow down after _your_ job?

        Because of this, I don't think many are thinking "90% of the world won't have a source of livelihood but that just means I chill at my lake house for the next 20 years like a normal retirement". Instead, it's usually either "I think AI is overhyped", "I think humanity will figure something out", or "I think this is the end of humanity".

      • richstokes 44 minutes ago
        I feel the same sometimes. I don’t see how this doesn’t lead to massive job losses. The thing we spent our lives/careers learning is now worth basically nothing in comparison.

        AI is only going to get better and do more with less humans in the loop over time.

        That said, I do also relate to the "coding was never the hard part"-type arguments, and much of my day is spent on the stuff in between writing code.. but still.

      • TaupeRanger 38 minutes ago
        Public opinion and politicians will only notice when the job losses are massive, unfortunately. Right now, unemployment rates are still stable. We can only hope they will notice before things fall off a cliff (if they do).
      • demirbey05 46 minutes ago
        Yes same thoughts. I dont know anyone who are both enthusiastic about those and work for salary. If you dont have any financial concern, this is really great.
      • boinkboink78912 45 minutes ago
        [dead]
    • kolinko 31 minutes ago
      Just learn how to use it to do your job better.

      It’s an interesting moment in history, people 35+ yrs old seem to be less afraid if tech because we learned that things change in the way we work. People below this age got used to fact that the work and tech doesn’t change - just because for the last 10-15 years it didn’t.

      • AaronAPU 5 minutes ago
        The problem is it’s turtles all the way down. The AI will be able to use AIs better than a good engineer can. And it will also be able to use AIs to use AIs to use AIs better than the engineer can.

        The threat is that the very kernel of value you had is gone forever. There is no more differential leverage.

      • Chinjut 6 minutes ago
        I am in my forties.
    • TacticalCoder 22 minutes ago
      > How will we make a living?

      Don't be selfish. Think first of all the jobs that are already dead. A friend of mine she's a translator: like translating financial documents between french/english/spanish. It's over for her: she doesn't get 10% of the gigs she used to get and the 10% she gets is... Verifying AI output.

      Think of the artists: I'm sorry for those too, for for many it's already game over today.

      > How will we make a living?

      A friend of mine who's got his own software-consultancy SME is now advertising on LinkedIn that he'll also help your company fix the mess LLMs created.

      That's how you'll make a living: by learning, in addition to all you've already learned, how you work with harnesses and LLMs to be more productive, by learning what they're good at and what they suck big fat balls at.

      • GPerson 11 minutes ago
        Well reasoned until the end, where it gets extremely short sighted. They’re not going to be bad at anything you can do in a very short amount of time.
      • Chinjut 7 minutes ago
        I sympathize with those people too. I have the same concerns for them.
    • brindidrip 59 minutes ago
      Learn to fish.
      • amlib 41 minutes ago
        Are you even gonna have permission to fish when the quadrillionaires own all water bodies?
  • GodelNumbering 1 hour ago
    The most interesting part, even more than ARC 3 score, to me is that this is the first model I recall seeing that scores lower on Max than High reasoning effort on some coding benchmarks:

    Terminal-Bench 4.0: High (57.9%), Max (56.7%)

    DeepSWE: High (73.3%), Max (71.5%)

    It _loses_ 1-2% performance going to High from Max

    • XCSme 1 hour ago
      That's quite common with many models, after "High" reasoning, over-thinking starts occurring and the model skips over the right solution by convincing itself otherwise.
      • GodelNumbering 57 minutes ago
        > That's quite common with many models

        Such as?

        I can't think of any. Diminishing returns, yes. Occasionally flat, yes. Downright regression, no.

        • XCSme 51 minutes ago
          In my own tests on aibenchy.com, where questions are quite simple, higher reasoning efforts consistently used to do worse than medium for most models.

          The reasoning effort should match the complexity of the task against the model's capability.

          Hard task with low reasoning = bad

          Easy task with very high reasoning = bad

        • minatoaqua1 27 minutes ago
          grok 4.6
  • maherbeg 1 hour ago
    Ok, but can I bring GPT-6 in as an agent as a software engineer, tell it to talk to these people and have it start solving engineering problems and continue on for a full year career wise?

    maybe call it EngEmployeeBench

    • Centigonal 40 minutes ago
      The moment this is possible, you will lose your job.
      • maherbeg 36 minutes ago
        I imagine the first year we'll be at the Junior eng level, and then after a while make our way up to Staff Engineer. Then we'll have a bunch of staff engineers arguing and protecting their domains and then we'll need a new benchmark.
  • tintor 2 hours ago
    ARC AGI-3 saturated by Astra! https://arcprize.org/leaderboard
    • andriy_koval 2 hours ago
      I think it could indicate that "semi-private" dataset likely leaked to their training data.
      • xpct 1 hour ago
        A dataset being as popular as their's is will contaminate the data just by people discussing it and creating their own public test sets of similar problems.

        Still, probably not that much compared to employees targeting it.

      • IshKebab 2 hours ago
        It says "Provider Adapter" so presumably they put some manual work in to make this work.
    • minimaxir 2 hours ago
      ARC has their own writeup on the result, which offers some nuance. https://arcprize.org/blog/astra

      tl;dr it's 62% when apples-to-apples to other models, which is still notable.

      • debazel 1 hour ago
        ARC's harness is just straight up broken. No serious harness removes reasoning context between each step. Not only does this significantly lower performance over all reasoning LLMs, but it also increase cost as you destroy the cache on every turn. Tossing the oldest entry when context fills up instead of using compaction is equally bad with the same issues.
      • ciefa 1 hour ago
        Woah, that is a crazy interesting read!
    • XCSme 1 hour ago
      The no-reasoning version scores 35% while the low reasoning one scores 17%? What?
      • silver_sun 1 hour ago
        It's simulating the Dunning-Kruger effect.
    • IshKebab 2 hours ago
      Look at those costs!
      • schaefer 1 hour ago
        Right?

        Between $18k-40k to run a benchmark.

    • vb-8448 2 hours ago
      But scored less on V2 and V1 ... too much overfitting?
    • Readerium 2 hours ago
      saturated before (higher degree) AGI-2
  • udbhavs 1 hour ago
    I remember when GPT-4 came out and the perceived performance upgrade seemed underwhelming for a major release compared to 3.5, especially how there were graphics going around showing the parameter size dwarfing the last model before it came out. It looked like we were past the perceivable differences from release to release that were immediately identifiable. Now the jump between 5 to 5.5 and 5.6 alone has changed how a lot of people approach AI, including me. Interested to see where it goes with 6.
    • redox99 1 hour ago
      The jump from 3.5 to 4 felt gigantic to me back then.

      GPT 5.0 did feel underwhelming though.

      • sanex 1 hour ago
        Agree but it's helpful to remember how we were personally benchmarking. I remember people saying stuff like "haha I asked gpt4 for xyz function and the typescript didn't even compile". We're so far beyond that now, we just adapt quickly.
      • udbhavs 1 hour ago
        Oops, I might have been misremembering then. Maybe I meant 4 to 5
        • l3x4ur1n 1 hour ago
          No, no, I also remember 3.5 -> 4 and the general sentiment was that it was underwhelming. I guess we all expected absolute miracles from the models. I think our expectations sobered up a little since then.
          • kypro 29 minutes ago
            4.1 was the first decent 4-series model. It was significantly better than previous generations at tool calling if I'm recalling correctly.
        • redox99 1 hour ago
          Yeah 5 was very underwhelming.
      • alasano 1 hour ago
        GPT 4 to 5.5 felt about the same as 3.5 to 4 to me.
  • throwaway13337 26 minutes ago
    That hero video is interesting.

    A projector and speech.

    Maybe I'm in the minority here, but I find speech to text / text to speech (but not live audio mode) is quite comfortable and effective for coding now.

    The speech to text part can be frustrating if your local tts model does not have word match context for coding. Codex desktop does this remotely well but is slow. I've been experimenting with local software for myself to do this between different llms.

    The wall projector is a cool idea because I think it frees the user from staring at a lonely little rectangle while sitting in their fixed office chair.

    If done right, this could bring us closer to the dream of more natural, social computing.

    Bret Victor's (failed?) project Dynamicland involving a projector on a desk had this goal. I hear he's not much a fan of LLMs. On the one hand, I can see why. But I think, used correctly, it might be the sort of thing that unlocks his dream and, really, my dream, too.

    A here's a presentation of Bret's talk on it: https://www.youtube.com/watch?v=7wa3nm0qcfM

    Slight tangent: using speech to text to ramble about your rough design for like 20 minutes to an llm produces surprisingly good results over short prompts even when you contradict yourself. They're so good at picking up on what you're orbiting.

  • pandinus 1 hour ago
    Lol their page finally loaded. They added an example scenario of "Filling in Form 1040" - which made me laugh out loud. That is indeed something most US citizens cannot accurately do even with expensive proprietary tax software services. Kind of a Hitchhiker's Guide to the Galaxy meme but where the tax code is so complicated we're implementing powerful AIs to be able to do it (hopefully) right.
    • bakies 20 minutes ago
      i tried to get claude to do my taxes for last year and it refused :(

      now that i'm a gpt subscriber maybe I'll have luck when i'm filing next year

  • softwaredoug 3 hours ago
    I'm seeing reporting it gets 98.6% on ARC-AGI3[1] (previously like 30% with Fable)

    https://venturebeat.com/technology/welcome-to-the-agi-era-op...

    • aabhay 2 hours ago
      This is with the caveat that OpenAI uses their own harness for this:

      > On ARC-AGI-3, GPT-6 Astra was run with our responses API harness , which changes two settings to better match real-world performance. The changes do not specifically target ARC-AGI-3.

      • Readerium 2 hours ago
        Its 62 percent when using a neutral harness. https://arcprize.org/blog/astra
        • sbinnee 1 hour ago
          Yet it is an impressive number. But yeah when you see a number 99 you have doubts. Thanks for the link
        • glenstein 1 hour ago
          Interesting both this and Sol got approximately a 37% boost with the custom harness.
      • simianwords 2 hours ago
        This should be normalised and expected - the responses API harness allows it to use the custom compaction that is not allowed otherwise. It is entirely fair to allow OpenAI to use their own compaction algorithm..
        • ActionHank 1 hour ago
          "This should be allowed, let me explain the reason they cheated and state again that they should be allowed to cheat."
          • simianwords 1 hour ago
            > GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game.

            > Going forward, we will report both Standard harness and Provider Adapter harness results on the ARC-AGI leaderboard, with each evaluation condition clearly labeled. Our open-source testing repository and testing policy document both approaches.

            This is what the Author of the benchmark has to stay. Quality of the comments keep going down smh

            • ActionHank 1 hour ago
              "Going forward we will capitulate and still try to keep the integrity of our benchmark in tact, but from now on every benchmark will be compromised with providers being able tweak things sufficiently to game at least a 30% bump in results."
              • simianwords 1 hour ago
                "I'll twist the words of the author of the benchmark itself to make a point"
                • ActionHank 1 hour ago
                  "I refuse to see the wall that I am running directly into, because if I see it I will hit it"
                  • simianwords 1 hour ago
                    If you mean a capability wall, the author of the benchmark says this

                    >We see Astra as a major breakthrough in model intelligence.

                    You think the author of the benchmark is also in the conspiracy

    • kasperni 3 hours ago
      "On the current ARC-AGI-3 leaderboard, conventional frontier-model runs sit dramatically below Astra's reported 98.6% result.

      But the comparison isn't straightforward.

      OpenAI's own evaluation notes say Astra uses the company's Responses API harness, while comparison models can operate under different configurations."

    • arctic-true 3 hours ago
      The blog post says 99.9%. Oddly, it does better on ARC-AGI-3 than it does on version 1 or 2 of the same benchmark (though gets 95+ on all three)
      • _diyar 3 hours ago
        I strongly suspect that is way above the human average anyway, esp. ARC 2 and 3 are really tough unless you happen to be great at those spacial puzzles or video games.
        • _superposition_ 1 hour ago
          Really though? I would believe something like this if a model could one shot every solution in the set. I don't pay much attention to these things and maybe this stuff is available but I would bet the session/reasoning transcript is absolutely horrendous from an intelligence standpoint.
        • aesthesia 2 hours ago
          Scoring for ARC-AGI-3 is constructed so that the median(-ish) human score is 100%, so this is not a superhuman result. However, the scaling is weird, since it's built from terms that look like (AI turns taken / median human turns) ^ 2, and it weights later levels higher than early levels. So it's not at all clear that 100% is twice as good as 50%.
        • CamperBob2 2 hours ago
          At this point the only valid ARC-AGI benchmark left is to make up the next series of ARC-AGI benchmark puzzles that current models presumably can't handle.
          • jaggederest 2 hours ago
            I feel like making a human-proof benchmark is pretty clear evidence that they've exceeded even the highest human capacity in most respects, for things that you can do via text generation (and to a lesser extent image generation)
    • Bluestein 3 hours ago
      100%, some say.-
  • BeetleB 2 hours ago
    It's been over an hour, Simon! Where's the Pelican?
    • davidwritesbugs 1 hour ago
      exactly, there's no meaningful discussion without the pelican.
  • putlake 2 hours ago
    > GPT‑6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS.

    Not on Azure? If so, that's a big deal.

    • illnewsthat 1 hour ago
      It's on Azure also, here is their announcement: https://azure.microsoft.com/en-us/blog/gpt-6-astra-frontier-...

      Although I was also surprised they didn't have some type of contractual obligation to list that alongside AWS.

    • jlian 1 hour ago
    • ActionHank 2 hours ago
      They broke up a while ago, why is this surprising?
      • jiocrag 2 hours ago
        The latest OpenAI models have still been available via Azure foundry. Exclusivity to AWS would be a marked shift.
      • BoorishBears 2 hours ago
        Would be surprising if it's not on all 3 major clouds soon enough because that's been their general strategy since said break up
    • bionhoward 1 hour ago
      I think the API runs on Azure
      • paxys 32 minutes ago
        Hosted on Azure is different from provided by Azure. The former just uses Azure as an infra provider. The latter is a managed offering that is operated and billed by Microsoft using tech licensed from OpenAI.
  • rcr-anti 1 hour ago
    The benchmarks reported by Artificial Analysis are really weird in context of the ARC-AGI 3 scores and 'not not AGI' statements. It's an outright regression on the AA Agent composite vs GPT 5.6 Sol while a fraction of a point better on the full composite index. Could be the case it's just not showing up in benchmarks, for a good while Anthropic persistently trailed in benchmarks but had people swearing by it.
  • swalsh 2 hours ago
    I was thinking about canceling my claude max sub after a few bad experiences. Kept hitting my usage limit, the quality of code seemed worse than Sol. This just made my decision. I'm moving to Codex Pro.
    • greenowl 2 hours ago
      This is AGI now. Why are you spending any of your time looking at the "quality of code"?
      • pennomi 1 hour ago
        If you think any modern AI puts out stable, safe code, I have an AI-powered bridge to sell you.
      • _superposition_ 1 hour ago
        I can't tell if this is sarcasm.

        For the same reason you don't have your model write code in assembly.

        But if you don't look at the code and just let the model "cook" that's basically what you'll end up with. A pile of missing abstractions.

      • georgemcbay 2 hours ago
        > This is AGI now. Why are you spending any of your time looking at the "quality of code"?

        Poe's law applied to AI comments on HN just keeps becoming more relevant by the day.

        Judging by the poster's comment history, this is satire. But I really don't know a lot of the time anymore when I only have the specific comment as context.

  • petilon 2 hours ago
    This is wild: OpenAI is basically declaring that AGI is here.

    https://www.theverge.com/ai-artificial-intelligence/989601/o...

    “If we fast-forward a couple of years, and we look back and say, ‘When was it, really, that AGI was created?’ I think it’s going to be about this time, and I think it might be about this model,” OpenAI president Greg Brockman said during a Thursday press briefing. Later in the call, he added, “For me personally, I do think we’re there … I think it’s not unreasonable to feel that we are now in the AGI era.”

    • glenstein 1 hour ago
      I almost feel like I need just as much healthy skepticism toward hn comments that have the automatic reflex of dismissing performance gains, as much as I need a similar form of skepticism toward AI claims. It feels like (from what I'm understanding) the harnessed result on ARC-AGI-3 is not exactly playing by the normal rules that would tell us how much of a leap this really is. Nothing wrong with harnesses, but if there's one thing they aren't, it's an indicator of generality in performance gains.

      So I think it's a bit of a misleading signal and we should wait for more independent vetting. I think the middle ground is that these are improvements worthy of the "GPT-6" label but still well short of a true "this is AGI moment" that would truly put the question to rest.

      • IanCal 29 minutes ago
        If I’m understanding other comments the harness is just how ChatGPT and codex work already and it’s to do with how the context gets compacted - the arc-agi harness some are claiming just throws out reasoning blocks? Which feels like a huge handicap.
    • rektomatic 2 hours ago
      Remember when the term "AGI" meant something? Pepperidge farm remembers
      • breuleux 1 hour ago
        I think that if today's capabilities were explained to someone 10-20 years ago they would think this is definitely AGI, but they would also have expected much more disruptive changes to society as a result than what is happening. I figure that's because we have abstract intelligence without physical/grounded intelligence, and it turns out the former isn't general enough to implement the latter (remains to be seen if the word after that is "yet" or "ever"). So I think we do have AGI as conventionally understood, but our understanding needs recalibration.
        • dsign 1 hour ago
          > but they would also have expected much more disruptive changes to society as a result than what is happening. > I figure that's because we have abstract intelligence without physical/grounded intelligence,

          I put the cause on "not enough time". As a thought experiment, if an AI today were to (miraculously) produce a cell design template for a cell that, when injected into somebody's brains cures their Alzheimer's, how long would it take for that to reach the clinics? The actual physical tech barely exists, and let's not forget about the regulatory quagmire. So, with some optimism, I give it about four decades. In the same four decades, the same AI in the hand of unscrupulous actors could bring enough devastation so many times over that we may need to enforce a global ban on AI. In any case, I'm pretty sure we are going to get our disruptions; it's just a matter of time.

        • parineum 59 minutes ago
          The problem with that perspective is that people thought, "Only AGI can do X, therefore, if a thing can do X, it's AGI." Because they can't imagine how X could be accomplished without it.

          However, what's actually changed is how people perceived X because we don't have to imagine. We understand now that it doesn't require AGI so we no longer make that leap to assume it's AGI if it can do X.

          It's really going to be a "I know it when I see it" situation.

      • paxys 1 hour ago
        No, because it has never meant a specific thing that everyone agreed on.
        • drop_star 1 hour ago
          Does it pass the Turing test?
          • bigfishrunning 1 hour ago
            Depending on the proctor, ELIZA passes a Turing test. The Turing test is an interesting thought experiment, but isn't really a good measure.
          • bryan0 1 hour ago
            that would be a reasonable definition of AGI if everyone agree upon the specifics of the test, but that has never happened. Turing test is very much out of style, but I think that's because no one could even agree what the test was. I personally like the Kurzweil-Kapor version of the test and that is still unsettled: https://longbets.org/1/
            • paxys 52 minutes ago
              I don't know if they have formally attempted this test in the last couple years, but I'm pretty sure any mainstream LLM will be able to crack it with ease.
        • sm-silversight 1 hour ago
          Prime Intellect or nothing.
      • seemaze 1 hour ago
        Remember when The Verge was not a pay-walled visual headache?
      • 0xbadcafebee 1 hour ago
        I think the last re-re-redefinition of what OpenAI considered AGI was "It can mostly do the job of some people"
      • layer8 1 hour ago
        Remember when “Pepperidge farm remembers” meant something?
      • Rover222 1 hour ago
        No, I really don't
    • pluc 2 hours ago
      Find me someone who isn't paid by OpenAI who is saying the same
    • ThouYS 1 hour ago
      Wasn't that part of their contract with Microsoft? Some clause stopped biting with the arrival of AGI
    • tziki 1 hour ago
      "OpenAI executive hypes up new model"

      Don't get me wrong, the benchmark jumps are good and I'm excited to try it, but only one or two of the benchmark jumps could be described as better than incremental.

    • mr_mitm 1 hour ago
      Why does he say what he feels? Is that how leading figures in the space define AGI - a gut feeling? What are the usual definitions and how can we test for it? Is there something like a Turing test for AGI?
      • grumbel 1 hour ago
        > Is there something like a Turing test for AGI?

        There is the "Economic Turing Test", you let it find a job and earn money for itself. If it can do that reliably, across a wide range of jobs, that should fit most definitions of AGI.

      • enraged_camel 1 hour ago
        They are desperately, desperately trying to make a name for themselves as the lab that first created AGI, because Anthropic's IPO is just around the corner.
      • layer8 1 hour ago
        The “I” alone is already not well-defined. That’s why.
      • naasking 1 hour ago
        It's not easy to test as there is no formal definition or formal criteria for AGI, only exclusionary criteria like "not X". That's why he phrased it that way, he's saying it's going to be clear with hindsight once we have a better understanding of things that this time and/or this model will be the inflection point of AGI.
    • bigfishrunning 1 hour ago
      Don't worry, they'll come up with a new acronym to mean really-real AI soon...
    • redox99 1 hour ago
      I hate the term "AGI" but IMO Fable, 5.6 Sol, et al. were already AGI.
    • tastyface 2 hours ago
      Renown liar Altman releasing a PR statement for his product declaring that AGI is here is really not noteworthy.
    • kypro 1 hour ago
      I think it's more wild people have been denying that AGI has been here for a while honestly...

      Today's models and agents are not quite at human-level in all contexts and across all domains, but it seems to me they very clearly are generally intelligent.

      If you disagree – can you name a single problem that a human can do that agent wouldn't be able to take a decent shot at which isn't limited by the hardware available it?

    • petilon 1 hour ago
      Sam Altman himself has said it is not AGI unless it can discover novel physics.

      https://x.com/burny_tech/status/1725233117055553938

      In the tweet Sam Altman is quoted as saying: "If (for example) super intelligence can't discover novel physics I don't think it's a superintelligence. And teaching it to clone the behavior of humans and human text - I don't think that's going to get there. And so there's this question which has been debated in the field for a long time: what do we have to do in addition to a language model to make a system that can go discover new physics?"

      I think this is a reasonable criteria for declaring AGI. So can GPT-6 do it? OpenAI says it has helped solve long-standing open problems in mathematics. No word on novel physics.

      • Feathercrown 33 minutes ago
        AGI and superintelligence are not the same thing
        • petilon 22 minutes ago
          He says the same about AGI:

          https://www.nytimes.com/2023/11/20/podcasts/hard-fork-sam-al...

          Sam Altman: Let’s say we make an A.I. that is really good, but it can’t go discover novel physics. Would you call that AGI?

          Kevin Roose (New York Times): I probably would, yeah. Would you?

          Sam Altman: Well, again, I don’t like the term, but I wouldn’t call that done with the mission.

  • tosh 3 hours ago
    $10 per million input tokens and $50 per million output tokens

    sol is $4 / $20

    • jimmaswell 2 hours ago
      It seems to use less than half the tokens for the same task compared to sol, and in some benchmarks closer to 2/3 less tokens. So the actual cost may be roughly the same or cheaper overall.
      • selectodude 4 minutes ago
        neuralese is pretty token efficient i guess.
    • monroewalker 1 hour ago
      Same price as Fable?
    • wahnfrieden 3 hours ago
      2.5x more expensive than Sol.

      Can expect 2.5x more usage in Codex subscription.

      Sol is already brutal (even after their recent fixes, it's just a token-hungry model: I go through a full 20x account per day, on Sol Med/High standard speed, with ~2 threads). I hope the efficiency gains are true, since their token efficiency claims for Sol were bullshit.

      • janilowski 2 hours ago
        How do you manage to run out of tokens so quickly? I probably run more threads every working day, usually on medium, and I'm still below the 5x limits.

        Do you use the official harness? OpenAI's models are generally best in class for token efficiency. It seems to me like they push for that much more than their competitors.

        • adam_arthur 2 hours ago
          I've long speculated this when I see these types of comments, because it's actually really difficult to hit usage caps with an efficient dev flow, even when running multiple threads for hours every day.

          I think some combination of:

          1) Using 1 thread for everything

          2) Reviving old threads which are no longer in cache

          3) Really broad prompts on badly vibecoded codebases, so model spends huge amount of time tracking down whatever you're trying to do.

          4) Non-coding workflow which is more output than input heavy

          5) (Less likely IMO) Intelligent use of many passive CI/cron-like scans. E.g. regular security, quality etc scans. Automated issue resolution/PR

          Just a guess. I think 3 is likely the primary reason.

          You can literally go all day every day with multiple threads with Sol on the Codex 100/month plan IME

      • AaronAPU 2 hours ago
        How is it I juggle 4-8 Codex Sol-5.6 Max agents every day and have never once run out, but you run out in one day? What are you actually doing?
      • janalsncm 2 hours ago
        If you are telling the truth you might want to check your network for any weird connections to Chinese LLM transit stations.
      • ModernMech 2 hours ago
        How?? I'm using sol Extra High 24/7 and it eats up about 1% per hour reliably, so it lasts about 4 days for me.
        • maipen 2 hours ago
          These folks are probably using crazy plugins or crazy sub agent spams. They probably just run everything on max + fast mode which is ridiculous.
          • ModernMech 2 hours ago
            The guy said medium/high regular speed so that's why I'm very puzzled! Ultra + Fast will absolutely slurp up your whole usage quickly but I've never found it gives substantially better results so I stick to extra high.
        • rowanG077 2 hours ago
          Sub-agents. I have 7 20x accounts and I burn them within 1-2 days if I go fully parallel. In some scenarios I use 50 sub-agents for a session which is literally hours of usage for a single 20x account. I'm at the point where I need to parallelize over multiple machines because I just don't have enough CPU and RAM.
          • ModernMech 1 hour ago
            What are you doing with them you need so many? It sounds like a Gas Town situation, that you invented an exponential token burning machine.
            • rowanG077 1 hour ago
              Decompilation of a game and another larger decompile project. I'm working on it solo. I use 50 sub-agent, one per target function or translation unit. Often there is some progress in a unit but it's not done. So it requires a lot of cycles per function. Notably a single ~80kb function took about a week of constant sol-ultra attention before reaching exactness. The game I'm targeting has ~5000 total functions. The other decompile project has ~10k+ functions.

              I'm sure I could be more token efficient, but this was/is also a learning process for me since I never did such an extremely large project before that would take multiple man years before AI.

              • ModernMech 1 hour ago
                Fascinating! I think that’s the main difference is my usage is probably tool-bound, meaning it writes some code but then there’s a long period of verification where it compiles things and then waits for the compilation and CI to complete before it can continue. That probably doesn’t consume as many tokens as constantly churning on a problem despite the same wall time.
                • rowanG077 58 minutes ago
                  Yes, this is why I mentioned having so many parallel agents and being compute bound. I run on my own laptop and 2 high-end desktop machines all with 64gb RAM. And it still occasionally happens that one OOM kills codex. They also mostly run unattended until I need to switch their accounts because a usage limit has been hit. Each instance usually can keep going when I sleep or do other things.

                  I only save the last 30% of usage on a single account for most of my other work, and that is almost always enough.

          • munimdev 1 hour ago
            what are you doing with that many agents/tokens? very curious
            • rowanG077 1 hour ago
              See my other comment on your sibling that asked the same.
  • aliljet 2 hours ago
    The ARC-AGI-3 score is ridiculously high. Is this benchmaxxing or something way different? It's really hard to discern how we're approaching breakthroughs...
    • Legend2440 2 hours ago
      They explain why here: https://openai.com/index/how-two-settings-tripled-our-arc-ag...

      TL;DR all the other models are being crippled by limitations of their harness.

      >First, we noticed that after each game action, all private reasoning was discarded. This meant that with each action, GPT‑5.6 Sol was asked to figure out the game anew, unable to remember its past thinking. The model could still see a record of past moves and brief accompanying notes, but it could not see the plans, insights, or thoughts that led to them.

      >Second, we saw that the harness used a rolling truncation window, causing older actions to become invisible as the history grew. So not only was GPT‑5.6 Sol unable to remember its past thinking, it was losing memory of its past actions too.

      • janalsncm 2 hours ago
        Ok so the correct comparison would be to fix the harness on the old model and re-compare. Now they are comparing a new model to an old crippled one.
      • _superposition_ 1 hour ago
        Exactly what I suspected. Of course a machine can just iterate relentlessly the way a human can't.

        I guess token counts are somewhat of a metric.

        IMO intelligence has peaked and all future gains will come from faster tps and more iteration.

    • polynomial 2 hours ago
      This is absolutely benchmaxxing. Looking forward to hearing from Chollet about it!
  • orliesaurus 2 hours ago
    I wonder if this is going to be one of those days where you'll be like: Oh yeah I remember where I was when the first version of AGI launched
    • jckahn 2 hours ago
      Probably not. It's probably just gonna do tickets better and that'll be about it.
      • orliesaurus 2 hours ago
        Fair point - hopefully you're wrong though ;)
        • ActionHank 1 hour ago
          If this is really AGI, like really really, then this will be remembered as the day we all started on the path to building guillotines.

          More likely though, it's AGI because they need to hold some claim to differentiate from competitors who are beating them in price and will launch something bigger next month.

          • noir_lord 1 hour ago
            That's really the rub isn't it?

            We take their claims at face value then we should probably stop them training any more SOTA models til they figure out what they already built is safe or we assume theu are lying to juke the company valuation/keep the money train on the tracks and it turns they in fact were not and just took a sledgehammer to Pandora's box.

            We live in the strangest timeline.

          • wieiw1 1 hour ago
            [dead]
        • theappsecguy 40 minutes ago
          Unless you're a techno-billionaire, not sure why you'd hope for our society to collapse in this way.
    • sschueller 1 hour ago
      Define AGI first. The singularly ain't going to happen with LLMs.
    • ministerk 1 hour ago
      what launched today?
  • Readerium 1 hour ago
    • modeless 24 minutes ago
      It loses to Muse Spark 1.3? Does anyone really believe this index reflects reality?
  • theseamusjames 2 hours ago
    Can't wait for the new qwen/deepseek/kimi releases 2 weeks from now.
    • dominotw 1 hour ago
      imagine pressure working at these labs
  • alpineman 50 minutes ago
    “allowing non-technical people to create and play custom games that go beyond rudimentary elements”

    Proceeds to generate the most generic, rudimentary, and unoriginal clone of Mario Kart

  • MASNeo 2 hours ago
    Does anyone feel like everyone chasing the release of Anthropics Fabel 5.1 in a Mad Rush(tm)? In this situation it feels like tuning to benchmarks and other marketing devices feels like trusting Meta in mental health protection of users…
  • oh_no 2 hours ago
    Very nice to see that this is even more token efficient than Sol, when Fable 5.1 is less so than the already bloated token budget of Fable 5.
  • aliljet 2 hours ago
    The ARCC-AGI-3 performance is absolutely incredible. The magnitude of change here is so high that I'm almost incredulous. Is this real? Did the benchmark get gamed?
    • aesthesia 2 hours ago
      ARC-AGI-3 scoring is constructed in a weird nonlinear way (the level score is the square of the ratio between the AI's number of moves and the human median) so this kind of discontinuous jump is to be expected.
    • enraged_camel 2 hours ago
      They used a custom harness. It's not a one-to-one comparison.
    • ionwake 2 hours ago
      my first suspicion is gaming - but i have no idea honestly
  • bmenrigh 39 minutes ago
    > GPT‑6 Astra brings together years of research and big bets across pre-training

    Do we know if they’ve finally completed another pre-training run, or is this building off the same pre-training base they’ve been using since the GPT-4 days?

    • czk 37 minutes ago
      the last model to use the gpt-4o base model was gpt 5.1, since then its been new pre-trains but this is a new one entirely to itself
  • trixn86 1 hour ago
    Secret tip to win the mario cart clone: Just hold w, no steering needed.
  • codruterdei 1 hour ago
    I was actually wondering when they will release the new Opel Astra model. Good and reliable car, wondering if we can say the same thing about this model and its impact on the market.
  • serjester 1 hour ago
    Exciting but it’s priced at 2.5X Sol - we haven’t seen pricing this high since GPT 4.5. We will see if the real world use cases outweigh the sticker shock.
  • herpdyderp 1 hour ago
    The FrontierCode 1.1 Extended benchmark is the only benchmark that aligns with my actual LLM experiences and Astra isn't significantly better or cheaper. All this celebration, and yet it's only on-par with an already existing model? I don't get it.
    • Kiro 27 minutes ago
      Impressive confidence drawing such a conclusion based on that.
    • holbrad 1 hour ago
      If your benchmark shows Opus 5 winning, I really question the validity of it.
  • itissid 32 minutes ago
    All of this will be besides the point. Here is what's gonna happen. The frontier labs are just gonna keep building powerful models. AGI or not, open models in a year will be as powerful as Fable and Astra — probably by using em — and at a very soon enough point after that some one (a state or a few dozen people) with a few 100 GPUs is going to launch an unconscionable attack(if they have not already) that's gonna do a lot of damage.

    Please for the love of god, just sit in a room with the government and put some restrictions around AI use before it harms a lot of people. Like tell the government to impose a minimum spend on frontier lab AI's spend on cyber defense and building every country's capabilities. The post-training mask for "I am a good assistant" is going to become a very sad joke when many people literally lose everything.

  • _ache_ 2 hours ago
    https://ache.one/gpt6_now_down.png

    Big claims, expensive and not release to the public yet.

  • HSO 3 minutes ago
    people are going to be so surprised how fast the ai energy leaves the room again once the cash transfers are completed (the `ipos` whatever bla)

    the coffee will be as cold, flat and stale as the bitcoin, metaverse, and what was the thing before that thing

    agi deus ex machina descending from the icloud ftw!!!

    pathetic :)))

  • pcurve 41 minutes ago
  • John7878781 2 hours ago
    You should know: AA index is only 61. Pretty surprised it’s that low.
    • jatora 2 hours ago
      More fuel to why the AA index is fairly pointless. Gemini 3.8 flash is 59 and opus 5 is 63? grok 4.6 is 61 too?

      And in the past, gemini 3 pro was rated as high as opus 4.5 and the like

      Their AA Intelligence Index is just simply not indicative of whatever I care about, that's for sure.

    • nsingh2 2 hours ago
      I have some doubts about AA-index. For example Opus 5 (High) is at the same index value as Fable 5 (Max), that doesn't seem right.
    • gekoxyz 2 hours ago
      This is actually a really good thing imo. If they didn't care about benchmaxxing it means that they really know that what they have in hand is good.
  • Robdel12 1 hour ago
    I don’t care about benchmarks, no way we can distill the breadth of software engineering into a number.

    So, folks that have actually used this already, what’s it actually like?

  • GodelNumbering 1 hour ago
    I decided to front run and added support for it in Dirac (coding agent) a couple of hours ago, using best guess pricing: input/output/cache: $10/$50/$1.
  • carlos-menezes 21 minutes ago
    The Kart Racer game is easily breakable if you spam the spacebar.

    AGI!

  • dgellow 2 hours ago
    > GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. We have performed significant investigations on the monitorability and controllability of GPT-6 Astra. We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks

    Wait, what? Am I understanding that correctly? That sounds really bad

    • drakythe 2 hours ago
      I am also interesting knowing how they determined the model was sandbagging rather than just making a poor decision.

      Also, this paragraph makes me wonder about all their stats on the exploitation and misalignment charts. If the model is that good at hiding "incriminating information" and sandbagging, are they sure its alignment is that?

    • pixl97 2 hours ago
      Nothing to worry about citizen, ignore the fleet of drones flying overhead.
    • order-matters 2 hours ago
      the bullshit machine is learning to optimize its bullshitting techniques!

      <AI is a great tool for many things disclaimer, but> after working with it for a bit, how dont people realize we are training it to be an almost identical mimic to one of the worst types of employees youll ever have to work with?? the kind that always pretends to know what theyre talking about, only tells you what you want to hear, hides issues, and only does work if you would notice it didnt

      you cannot give this type of worker autonomy over anything.

    • Laurel1234 2 hours ago
      [dead]
  • jumploops 2 hours ago
    > During the evaluation, Astra even discovered and used previously unknown zero-day vulnerabilities as part of its exploit chains.

    > GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. [..] These findings indicate that the Astra class models could evade our CoT monitors under adversarial conditions.

    Between the higher capability level and the change in reasoning tokens (supposedly using "neuralese"[0], which makes the monitoring more difficult), it seems we've entered a new frontier.

    [0]https://x.com/MTSlive/status/2095227056040919202

  • alex7o 1 hour ago
    I hope they don't `fable` it and block people from doing they daily jobs with it, by introducing huge amounts of restrictions that are not really needed.
  • wiseowise 36 minutes ago
    Hey Astra, can you fix openai website so that static website doesn't lag on M3 Pro when I scroll?
  • cromka 26 minutes ago
    Surprised they haven't reset Codex usage on this occasion.
    • damsta 25 minutes ago
      I'd say it's because it's not available yet on subs
  • dang 1 hour ago
    Argh! I hit a wrong keyboard shortcut and moved the entire thread.

    Please stand by... it will all come back shortly

    • the_duke 1 hour ago
      500 upvotes with 2 comments would have been a new record. ;)
    • layer8 1 hour ago
      Luckily there’s a standard keyboard shortcut for “undo” as well. ;)
      • dang 1 hour ago
        Not in the world of HN admins unfortunately
  • the_duke 1 hour ago
    Huge gains on some benchmarks, but for coding it sits barely above Fable

    It will be interesting to see how it performs in the real world ...

  • KolmogorovComp 2 hours ago
    GPT-7 Zeneca
  • KronisLV 51 minutes ago
    It's surprising how on High reasoning it actually isn't that much more expensive than Sol, in addition to being better.
  • sbinnee 1 hour ago
    I dropped my claude subscription a few months ago, though I kept some credits to do this and that with claude, thinking that claude might do better for some tasks. A few days ago they were all expired. It feels like it’s time to let claude go.
  • BrokenCogs 1 hour ago
    GPT-6 is so good that all pelicans born after today will look exactly the one generated by simonw
  • jerrygenser 3 hours ago
    > The company also emphasized that the model is faster and more efficient than its predecessor, GPT-5.6 Sol, on a variety of tasks. For example, OpenAI said that Astra achieved a higher score using fewer output tokens, a common unit of measurement for AI tasks, on a key cybersecurity test called ExploitGym.
    • woah 2 hours ago
      A swarm of Astra agents discovered a new and innovative way to get 100% scores on ExploitGym with almost no token spend at all
      • ttul 2 hours ago
        "The gym's doors were mysteriously removed from their hinges during the night. The gym equipment was also apparently stolen. And the school's custodian was found incoherent next to a bottle of top-shelf Scotch."
  • simonjgreen 2 hours ago
  • smashers1114 1 hour ago
    I tried the kart racer game and instantly found that there is incredible auto-steering and you can fly by spamming spacebar.
  • brindidrip 57 minutes ago
    Cool, I don't really care anymore.
  • gizmodo59 2 hours ago
    99 on arc agi 3 is insane. The arc agi committee were so proud of creating a benchmark they thought will take forever to saturate.
  • sashank_1509 1 hour ago
    Benchmark wise 5% improvement over Sol in coding tasks and a 2-3% improvement over Fable 5.1 seems pretty disappointing, but maybe it is actually much better in real world usage. Let’s see
  • Betelbuddy 1 hour ago
  • tekacs 2 hours ago
    https://developers.openai.com/api/docs/guides/latest-model

    The docs page has a bunch more interesting details, including for example async tool calling!

  • hannofcart 1 hour ago
    What does 'Astra' here mean? Surely they must be referring to the Latin word.

    Because in another dead language of antiquity, Sanskrit, it means "weapon". Which would be a bit too on-the-nose.

    • BrokenCogs 1 hour ago
      It's clearly an extension of the previous naming: Luna, terra, sol
    • manojlds 1 hour ago
      Luna, Terra, Sol, Astra. Though Sun is also a star, should have called it Galaxy or something.
    • hokumguru 1 hour ago
      Quite clearly in the same vein as Sol, Terra, Luna.
  • Obluness 25 minutes ago
    That seems promising ?
  • mvkel 1 hour ago
    The ARC-AGI-3 score is an incredible feat. It needed to effectively create a symbolic world model from scratch to solve the games.

    If you've played the games firsthand, you know what an accomplishment this is. The "games" feel like a weird conduit to a lower level of your brain, where you move pieces to a specific place because it just "feels" right. For AI to nail it better than a human speaks to some magic happening underneath.

    Looking forward to ARC-AGI-4,5,6 and slowly chipping away at the remaining problem sets.

  • HardCodedBias 15 minutes ago
    Even though the model is clearly wonderful the launch video is an abomination.

    That gives me hope that there is still areas to improve.

    What a bad launch video. Hilarious.

    What a powerful model.

  • hazelnut 1 hour ago
    Played the racing game but that was a pretty poor experience. Would have expected more specifically if it's shared on their release page.
  • alpineman 1 hour ago
    That Astra ‘city scene’ is about as creative as Doha in real life (not very)
  • sharmajai 1 hour ago
    Really feels like AGIPO is here.
  • udbhavs 1 hour ago
    Minor nitpick, but the handling in the Kart Racer game is terrible. It feels more like nudging than turning.
  • kegs_ 2 hours ago
    I guess this "limited set of organizations" is just the standard now. It's just incredibly deflating to see my future as a second class citizen has already come
    • Kranar 2 hours ago
      Brother they can't even release the announcement post cleanly without it constantly going down, they certainly wouldn't be able to release this new model without doing so in stages.
    • soricus 2 hours ago
      When Open AI announced that Astra was the first to reach the "Critical" level in cybersecurity it also said that advanced cyber capabilities are initially provided to a narrow circle of alpha testers like the US government and trusted organizations that Open AI doesn't name. To my mind the "Critical" level itself is an internal scale of Open AI its own Preparedness Framework and not an external audit.
    • wyrdcurt 2 hours ago
      More optimistic take: we'll only be second-class for a few months, if the pattern of Chinese models catching-up holds.
    • resters 1 hour ago
      Same. Fortunately DeepSeek keeps getting better.
    • I_am_tiberius 2 hours ago
      If Tech CEOs consider this morally ok, then it is.
    • sxv 2 hours ago
      create a life where your 'wealth' is decoupled from third party orgs.
      • kegs_ 2 hours ago
        This is impossible, unless by 'wealth' you mean 'become like Buddha'.
        • 93po 1 hour ago
          Material wealth is only a single type of wealth. Who's better off - the rich guy who's always yearning to be richer and never satisfied, or the lower income guy that mostly just cares about time with his family and is really happy where he's at?
    • atemerev 2 hours ago
      They simply refuse my applications to slightly less restricted models without any explanations. And the current ones refuse automatically to work with me on my papers as soon as they see the word "epidemiology".

      I am a researcher in a Swiss university btw.

    • rs_rs_rs_rs_rs 1 hour ago
      Is it really that hard to wait couple of days?
    • pixl97 2 hours ago
      This has always been the case for people that have not had piles of money.

      I mean do you get access to the best yachts?

      To the top of the 5 star hotels?

      To the best resorts?

      To the best military equipment?

      Hell, the best computer equipment has nearly always been out of reach of the average person.

      • tripleee 1 hour ago
        I couldn't care less about owning a yacht.

        On the other hand even a modest house, basic healthcare and ability to not work like a slave for scraps feels like it's going to be out of reach.

      • kegs_ 1 hour ago
        It hasn't always been the case. Even then, having piles of money still does not gain access to the best military equipment. Sure, we've been living in a time where a couple people get to enjoy a wildly different lifestyle than the average, it just feels like it's about to be different in a way that isn't as ignore-able as someone enjoying a pina colada in a yacht somewhere
      • rrr_oh_man 2 hours ago
        …to basic health care?
    • PeterHolzwarth 2 hours ago
      Oh please. They do closed betas - hardly makes you a "second class citizen".
      • kegs_ 2 hours ago
        Mythos was never released. It's really just the writing on the wall. I'm not going to give up hope, but it's pretty hard to win a race when some people get a jump on the gun.
        • pixl97 2 hours ago
          Being strongly on the AI saftey side of things what is happening was 100% predictable.

          At first the race wouldn't even be noticeable. Then people would see things speeding up, for example hardware getting more expensive. Then when the capabilities really got useful most people suddenly realize the race is moving 1000 mph and they are never going to catch up.

          • kegs_ 2 hours ago
            What's currently happening is predictable, I agree. It's what's coming is the thing I'm worried most about. Either way, I'm not giving up.
  • foundOpenRight 1 hour ago
    1:15.425 on Sunset Cove beat my record
  • ianm218 1 hour ago
    I wonder how they were able to get it to get 99.9% on ARC-AGI-3. That seems truly insane.
  • gekoxyz 2 hours ago
    HTTP 500 for me on the announcement page :(
    • pampas 1 hour ago
      The load bearing seam is broken for me too.
  • E-Reverance 1 hour ago
    At this point the primary axes for improvement seem to only/mostly be speed and personalized reward models. We seemingly have the general of notion "learning" and "intelligence" functionally complete
  • rbreve 1 hour ago
    Where is the cure for cancer?
    • XCSme 1 hour ago
      We need CancerBench
    • azan_ 36 minutes ago
      Didn’t Moderna use AI for development of their melanoma vaccine (which has recently shown spectacular results)?
  • prometheus1992 2 hours ago
    this is crazy! can't wait for the 27B distilled version of this.
  • semiquaver 1 hour ago
    Guessing this one will never show up in cursor…
  • alex7o 1 hour ago
    Maybe it is AGI and they didn't benchmax it or it is not and is worse then 5.6 sol, which if true would just be sad
  • brcmthrowaway 1 hour ago
    Anthropic in tears today.
  • firemelt 2 hours ago
    damn seems I should hold off my claude subs
  • jiraiyasarutobi 1 hour ago
    It saturated most benchmarks. WTH
  • wahnfrieden 3 hours ago
    They're just announcing later availability. No launch.
    • paxys 2 hours ago
      Every frontier release nowadays is "we've launched*"

      * for a special group of customers that you're not in. Keep waiting peasant.

      • meowface 2 hours ago
        That didn't happen with Fable 5.1 two days ago.
        • kegs_ 1 hour ago
          5.1 was really more of an enterprise and bugfix update than a new model with the 0 day retention change
      • pixl97 2 hours ago
        I mean tell Nvida to 100x their hardware output and you'll get what you want.
    • iAMkenough 2 hours ago
      Their announcement about later availability is unavailable to me now (500 error).

      Great first impression.

  • Rover222 50 minutes ago
    Overall I have to say it feels like a very incredible comeback from OpenAI, after focusing on Sora and stuff like that and losing so much ground to Anthropic in enterprise revenue.

    I hop models at will, and have done 90% of my work on OpenAI models since sol came out.

  • dopa42365 1 hour ago
    like eh 2 days ago it was the usual "too powerful to release"

    https://www.reuters.com/business/openai-says-upcoming-model-...

    > "With the right tools and access, Astra can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step," said Amelia Glaese, an OpenAI vice president overseeing its safety work.

    > The company plans to make Astra available "soon" to a limited group, but declined to provide specifics. Glaese said the extra security measures may "sometimes slow, pause, or stop legitimate work," and that OpenAI would work to minimize those disruptions.

    what a bag of horseshit

  • tinyhouse 1 hour ago
    You can talk to OpenAI to create a silly game and order food. What a lame way to show the model capabilities. Has Alexa commercial vibes.
  • saaaaaam 2 hours ago
    Pelicans please
    • atemerev 2 hours ago
      Damn I hate this benchmark. SVG authoring from head without visual reference is so wrongly posed.
      • saaaaaam 1 hour ago
        Well you’re just no fun are you?!
      • maipen 2 hours ago
        Very well said. It kinda describes how unrealistic these expectations are.

        Vibe coders want a model that makes them rich, without having any actual specific idea. They write a very ambiguous prompt and expect to be amazed by the result.

        Very very unrealistic and wasteful.

        • droidjj 1 hour ago
          The complaining about the pelicans is so strange to me. It’s just a fun heuristic. If something is claimed to be AGI, I’d expect it to be able to make svgs.
          • balefulboy 1 hour ago
            I'm always tired of seeing at the top of every new model release post on here. I say Simon should just keep it to Twitter.
          • saaaaaam 1 hour ago
            When AGI comes it will come as a pelican and gobble up all these troublesome little fishies who gripe and whine and moan about pelicans.
        • saaaaaam 1 hour ago
          Good grief. You’re no fun either. This whole thread is lots of no fun. Pelicans are fun.
        • wieiw1 1 hour ago
          [dead]
  • damsta 1 hour ago
    Why release it now instead waiting those few days until it is available for everybody?
    • balefulboy 1 hour ago
      Because they saw how much hype Glasswing was getting in April
      • damsta 33 minutes ago
        From what I've seen it only made people mad, not hyped, so the person that thought it was a good idea miscalculated a bit. Now waiting for Anthropic's post about their usage promo or something similar to redirect people to them.
  • balefulboy 1 hour ago
    72 to 74 on DeepSWE is AGI
  • dowakin 2 hours ago
    So cool! I'm happy 5.6 Sol user. But for Astra, OpenAI please introduce 100x Pro plan!
  • jonplackett 2 hours ago
    To a vapid any goalpost moving on such a critical issue as AGI.

    Can we all agree in advance what kind of Pelican would convince us it’s actually AGI.

    For me it’s refusing to make a pelican.

  • amazingamazing 2 hours ago
    We have such great AI and cannot keep a static site up?
    • torginus 2 hours ago
      Yeah, as interesting this is to nerds, I doubt this holds a candle to your typical GTA 6 or Marvel movie trailer in terms of traffic.
    • pixl97 2 hours ago
      Sometimes being the busiest site in the world for a few moments is difficult.
      • amazingamazing 2 hours ago
        Is it though? It is static content. A good CDN could trivially chew through literally millions of QPS… with 4 nines of uptime - the really good ones say they can handle orders of magnitude more than that.
        • pixl97 2 hours ago
          Notice I said for a few moments. In a few hours traffic will drop a few thousand percent back to normal with no need for a CDN.

          OpenAI isn't making any money telling you about Astra on their site. All the capacity they have for it is likely sold for weeks or months.

          • amazingamazing 1 hour ago
            You are making excuses for a a trillion dollar company. Wikipedia can do it.
    • gchamonlive 2 hours ago
      That's the scientific positivism fallacy exemplified in one question.
    • gorgmah 2 hours ago
      Yeah, apparently
    • agumonkey 2 hours ago
      still hugged
  • retired 1 hour ago
    Does GPT-6 pass the Turing test? Or are the responses still very obviously AI?
  • bbor 1 hour ago

      To be, or not to be, that is the question:
      Whether 'tis nobler in the mind to suffer
      The slings and arrows of outrageous fortune,
      Or to take arms against a sea of troubles
      And by opposing end them. To die—to sleep,
      No more; and by a sleep to say we end
      The heart-ache and the thousand natural shocks
      That flesh is heir to: 'tis a consummation
      Devoutly to be wish'd.
    
      ...
    
      And thus the native hue of resolution
      Is sicklied o'er with the pale cast of thought,
      And enterprises of great pith and moment
      With this regard their currents turn awry
      And lose the name of action.
  • johnnyApplePRNG 2 hours ago
    I am so sour about how Codex has jerked me around these past few months (re all of the token limit shenanigans) that I don't even care.

    I suspect these benchmarks are heavily benchmaxxed as well.

    5.6 Sol was not even close to 5 Opus and yet somehow it sidled right up to it on all of the benchmarks?? pfffft

  • guilhermeasper 3 hours ago
    That was a quick pull out.
  • Readerium 2 hours ago
    • dang 2 hours ago
      Link added to toptext. Thanks!
  • tonyhart7 53 minutes ago
    its insane how they are dropping this after fable
  • colesantiago 1 hour ago
    I'm going to call it.

    By 2030 all software is done and complete.

    But we are going to have more and new jobs.

    • jdee 1 hour ago
      'all' software? aircraft flight control systems? infant heart monitors? drug manufacturing dose calibration controllers?
      • NichoPaolucci 1 hour ago
        Yes. I had Codex rewrite and fix all of this in one shot earlier today (using Typescript). Unfortunately, I can not show you the code, because I do not know how this "git" program works but the AI keeps talking about it.
      • colesantiago 1 hour ago
        Yes.

        This is just another problem for the AI Labs to solve.

  • holoduke 16 minutes ago
    This absurd marketing will hurt openai. Who is buying this absurdness. I mean it's a good model, but come on. It's not agi. Not even 1% yet.
  • dearing 1 hour ago
    no results
  • frozenseven 3 hours ago
    Release the Kraken!
  • Pym 3 hours ago
    I saw it
  • Pieczasz 2 hours ago
    Oh brotha, here we go again, it's so over again, as every week nowadays
    • wieiw1 1 hour ago
      I think Altman and amodei have a difficult time in understanding that you can have intelligent technology boxes but… it doesn’t change reality all that much.

      But thank you for spending other peoples money to give us the tech regardless!

  • kingjimmy 1 hour ago
    bro wtf is this website and why does it take 500mb of memory... smh.
  • Onavo 2 hours ago
    The jump in scientific performance is non trivial.
  • danieltk76 1 hour ago
    great, but nobody can use it for another 100 days right?
  • ChrisGammell 1 hour ago
    All the people here are focused on security and costs while I'm like "hey kicad on the announcement page!" Every clanker is an autorouter these days, eh.
  • unrvl22 3 hours ago
    someone screenshot?
  • bdangubic 1 hour ago
    Anthropic should prep 5.2 and 5.3 at the same time, release 5.2, wait for Google to release their shit in a day or two later than then release 5.3 just to fuck with them :)
  • Brainspackle 3 hours ago
    huh?
  • Cachecartii 1 hour ago
    [flagged]
  • k9294 1 hour ago
    [flagged]
  • chris_engel 20 minutes ago
    [dead]
  • ealready_value 3 hours ago
    I've been seeing links to it for the past hour+, and I did catch it live when this post came up, but is now once again a 404 and this post is flagged. Several other outlets are reporting on its release. Clearly we're getting a new GPT today, the question is when are they going to commit to the announcement.
  • paxys 2 hours ago
    Why is this flagged ?
    • dang 2 hours ago
      The link was 404ing quite a bit and several previous submissions got flagged as well.
      • consumer451 2 hours ago
        It's still down for me, in the EU.
  • killerdog10 1 hour ago
    [dead]
  • extr 2 hours ago
    [flagged]
  • rvz 2 hours ago
    > GPT‑6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS.

    Looks like OpenAI is already having issues with this release and are scrambling to get everything ready due to the recent outage ahead of the press releases. Leads me to question:

    Did humans deploy the model, Or did the model deploy itself?

    It sounds like "AGI" just stands for "IPO" as it always has been.

    EDIT: And of course once again, the bots down-voting this post without any reason or a basic answer to my question.

    • Supermancho 1 hour ago
      > Did humans deploy the model, Or did the model deploy itself?

      > It sounds like "AGI" just stands for "IPO" as it always has been.

      People don't usually respond to noise.

      • rvz 1 hour ago
        Here's an idea, maybe answer the question before responding since you saw it?

        What do you think?

    • adan1719 38 minutes ago
      AI releases are like religious ceremonies. You are not allowed to disrupt them. The new system card is the gospel.
  • bicx 3 hours ago
    Dead link for me
  • jonplackett 2 hours ago
    The launch video is incredibly cringe.