44% on ARC-AGI-1 in 67 cents

(mvakde.github.io)

54 points | by porridgeraisin 1 hour ago

5 comments

  • xeonax 51 minutes ago
    Even cooler is his about me mention of saving his own life https://mvakde.github.io/ > Saved myself in a medical emergency (doctors didn't know what rhabdomyolysis was)
    • qlm 16 minutes ago
      Crazy, considering rhabdo isn't that rare.
      • p-e-w 8 minutes ago
        “What do you call a medical student who graduated at the bottom of their class?”

        “Doctor.”

  • pwmglenn 9 minutes ago
    Really impressive and creative research. I wonder if the leading labs do anything similar with their models? It doenst look like the open source labs do?
  • embedding-shape 52 minutes ago
    Is the author only running their model against one benchmark? I don't think anyone finds that difficult to achieve, the difficulty comes when you want to make the model not benchmaxxed to a specific benchmark, and generalize so it can solve problems not part of the training data, but seems this model is specifically for not this? How useful is that?

    If you just wanted to pass these specific tasks in this specific benchmark, and wanted to do so cheaply, I'm sure a non-LLM-based approach would yield better results for even cheaper, since what the author's model does, seem to basically be "solve ARC puzzles", not a general LLM or "coding" LLM.

    • bkaae 30 minutes ago
      I read this as a response to the current hype around LLMs. He is showing computers can solve these issues, without using an LLM architecture. A lot of people have sort of forgot that machine learning is more than just LLMs these days.

      I found it to be a very interesting angle.

      • embedding-shape 25 minutes ago
        > He is showing computers can solve these issues, without using an LLM architecture.

        Isn't it a LLM he's building though? My very point is that this particular use case could be solved better without building a LLM, now you claim he is not? The description of what he's doing surely makes it sound like it's a (very small) LLM, and personally I'm still on the "if it quacks like a duck" train in life.

        > A lot of people have sort of forgot that machine learning is more than just LLMs these days.

        Yeah, which I guess if you make my previous comment more concise, is exactly what I state too.

    • f311a 32 minutes ago
      The whole point of his model is to optimize for a very specific benchmark.

      BUT, he does not use labels when training, so the model does not know the answers.

      • embedding-shape 24 minutes ago
        > The whole point of his model is to optimize for a very specific benchmark.

        But benchmaxxing is what we generally try to avoid for training, as there is no point really for it. We used to call it "overfitting", now you're saying this person does it intentionally? Why?

        • K0balt 1 minute ago
          There are plenty of applications where a machine learning system needs to optimize for a very limited data set that is still intractable by linear logic systems of reasonable scale and complexity. It’s interesting, because he is using the legos of LLMs to build highly specialized machine learning systems, which is a very pragmatic approach. Obviously a lot of other ways to achieve similar goals, but it’s cool to see someone back porting the modern tools towards older style optimizations.

          Also, the complexity of the task he is using occupies an interesting middle ground of ultra high dimensionality (for a “simple” problem) while being limited in width to a narrow set of solves- a space where one would be tempted to imagine you would need a much more capable system.

        • f311a 18 minutes ago
          Why not? There is $700k reward for the next iteration of this benchmark https://www.kaggle.com/competitions/arc-prize-2026-arc-agi-2...

          I would not call this overfitting, it's finetuning for specific task where you have a benchmark.

  • larodi 13 minutes ago
    "I don’t understand why others didn’t figure this out"

    - how about we allot the possibility that so many of presumed ML experts don't have any clue what they be doing, and are eventually API bitches, nothing more.

  • eis 40 minutes ago
    > Increases in LLM scores are now mainly driven by post training (evidence in next section) and are probably a function of amount of synthetic data. They are learning to solve ARC tasks, not learn general abstract reasoning

    Agreed and that's for any benchmark. Private tests are better but you still have to trust the provider to not log and use them for training.

    That's why I like when a new set of tests like a new ARC-AGI version is published, that's where you can see which of the models abstracted to more general capabilities instead of being focused on the previous tasks. Most models completely fail new ARC-AGI tests.

    The "67 cents" part though is misleading imho. You can't extrapolate from there and think that investing say $100 will get you a lot better results. You hit a ceiling very fast and investing into more compute will give you diminishing results. So yes, you can train a custom model to do somewhat decently on a specific set of tasks but then what?

    • bkaae 28 minutes ago
      Then nothing - that's awesome. People think that LLMs are the know-all do-all solution to every problem now.

      Putting solutions in terms of cents is a great way to potentially win over some ai boosters imo. There are other ways to solve hard problems.