Kolibri Has Landed: A Sovereign Open-Weight Model

(aleph-alpha.com)

107 points | by bastitx 5 hours ago

8 comments

  • amoshebb 18 minutes ago
    Qwen3.8 27B beats Kolibri 79.9 vs 70.8 in German in Kolibri's harness on Kolibri's benchmark.

    Also, once the Cohere takeover is complete will they still be able to use this "sovereign" claim despite being 90% owned and 100% operated out of Toronto?

    • isusmelj 11 minutes ago
      I was also quite surprised to see that. Considering the effort Aleph Alpha put into to their new model, it seems like the Qwen team needs access to vast amounts of german data o.0 I'm really glad for these efforts for open models from within Europe.
  • kkm 57 minutes ago
    Thank you Aleph Alpha team for making it open.

    We as many other’s were curious to try and benchmark it.

    On that note, as a small gesture of support, we’ve hosted and made Kolibri-1 free for anyone to try for the next few days.

    No GPU. No setup. Just try it. tesseracted.com/kolibri-1-chat/

    https://x.com/konarkmodi/status/2106373678589960260?s=46

  • petesergeant 1 hour ago
    I wish nothing but luck for an EU model, but:

    > intellectual-property safety

    My suspicion is that you simply can't build an even slightly competitive model without liberally stealing your training data, in 2026, as much as I'd like it to be otherwise. You can get to the point that I suspect most of the frontier labs are at, where you've laundered the initially stolen data through the creation of huge amounts of derivative synthetic data, but still. Anyone who isn't comfortable stealing their training data is bringing a knife to a gun fight, and is going to die a noble but inevitable death.

    • rpdillon 59 minutes ago
      This doesn't seem to be true. There's a clear legal path via the first-sale doctrine to train models on copyrighted works. It's been years now, and publishers still don't seem to be offering anything for training (e.g. bulk licenses solely for training use), but adversarial interoperability via cutting up books and scanning them remains perfectly legal.

      There's also the ability to distill other models, which is also not illegal (though I'm sure they like to come after whomever for TOS violations, but thats a civil matter).

      And, of course, the obligatory copying-isn't-theft observation. A recent supreme court judgment put it well.

      > Since the statutorily defined property rights of a copyright holder have a character distinct from the possessory interest of the owner of simple “goods, wares, [or] merchandise,” interference with copyright does not easily equate with theft, conversion, or fraud. The infringer of a copyright does not assume physical control over the copyright, nor wholly deprive its owner of its use. Infringement implicates a more complex set of property interests than does run-of-the-mill theft, conversion, or fraud.

      Folks are pretty smart here, I think we can handle these nuances, even if we don't agree about whether they are good.

      Edit: reading through the full text of their post, it looks like they are using common crawl, which is likely just as much of a copyright infringement as Anna's Archive -- it's not like published works have a unique claim to copyright. I think this strengthens your point, though: I was expecting to see scans as training data, but it doesn't appear to be the case.

    • torginus 29 minutes ago
      hasn't IP law passed the statute of limitations? As in most models are probably trained on output of other models, as creating enough data otherwise is not feasible. Additionally, they are trained on github repos made since the AI boom, which were generated by models with IP issues (who knows what and how).

      Thus training on 'clean' data is like trying to unscramble an egg.

    • ekidd 53 minutes ago
      This model isn't terrible, at least on the benchmarks. It's 78B A3B and performs about like Qwen3.6 35B A3B. You can probably run it comfortably in 96B of RAM with a decent quant that doesn't lose too much.

      Unfortunately, Qwen3.6 35B A3B isn't really a useful coding model. You'd probably want Qwen3.8 27B at a minimum, which requires at least 32GB of VRAM (not system RAM) to run semi-comfortably.

      So this isn't going to be a competitive model for hobbyists, and you'd have to be a bit desperate to use it for coding. But if you work in a regulated industry and don't mind paying for a bit of extra hardware, it isn't catastrophically bad, either. Probably would work fine for information extraction or as a "classifier" like Jev. (Almost any GGUF model can be turned into a classifier using llama-server. See pi.dev codemode for sample code.)

      So they're not a real contender yet, but they look like they're probably at least minimally credible.

    • embedding-shape 54 minutes ago
      Have you tried the model itself and seen if it's "even slightly competitive" or not, and have specific complaints about it? Otherwise it feels like you're complaining about something that is easy to test but rather than taking the time to actually figuring that out first, you're arguing about some general and theoretical thing which the submission (may) directly disprove.
      • petesergeant 48 minutes ago
        No, I haven’t, but I’ll donate $20 to the non-political charity of your choice if it doesn’t turn out to sit a significant difference from the frontier.

        I think it’s a safe assumption that they’re leaning into “sovereign” because performance is bad.

        • ygjb 11 minutes ago
          I think you have it backwards. Sovereign is the goal, good can come later.

          There is a proliferation of sovereign models under development specifically to address data sovereignty, and a loss of performance is absolutely acceptable over the risk that a once ally will turn adversarial, or a foreign business stops serving what has become critical infrastructure.

    • Zambyte 57 minutes ago
      Is it noble? The entire notion that training data can be "stolen" at all is quite silly. If I "steal" content that someone created to use for training, what am I actually stealing? They didn't lose anything. They still have everything they had before. What was "stolen" was "unrealized profit", or put another way: money that wasn't theirs, that they had no entitlement to. The only actual crime that is committed is "unauthorized copying", not stealing. Support and enforcement of copyright feels wildly authoritarian. It's hard to see it as noble.
      • folkrav 31 minutes ago
        The same could be said of any digital product being sold. Nobody actually loses anything but the actual sale either when you download a cracked game or piece of software, a movie, music, etc.
      • pepperoni_pizza 46 minutes ago
        That's fair, but then people like you complain when someone "steals" I mean distills openai or anthropic models.
        • brookst 38 minutes ago
          It’s poor form to argue against someone by imagining something totally different that they might believe, which would then make them hypocritical.
        • Zambyte 44 minutes ago
          I don't complain about that. Model distillation is excellent.
  • tosh 1 hour ago
    i wonder if the custom tokenizer is better in practice, the examples look interesting though
  • oblio 30 minutes ago
    I wonder if we can start having LLM distros: community led distributed training runs with periodic releases, open weights, FOSS code, the whole shebang. Maybe the public training sets can reach a level where an LLM trained on them can be good enough for most things, such as web search and aggregation, coding, etc.

    I wonder how far we are from this. How far are we from LLM's Debian moment?

  • mistyvales 1 hour ago
    The Sega 32X game??
  • cyanydeez 2 hours ago
    Interesting they recommended high end software without considering quant 4 or 8 and still used A3B which should give good throughput on cheap hardware.

    If they can follow Qwen3.8-Flash-Next, the could draft off the huge reduction in VRAM requirements.

  • OmerKaraaslan 36 minutes ago
    [flagged]