Aleph Alpha releases open-weight Kolibri with 1M context
Kolibri is being released as a downloadable Apache 2.0 model for on-premises sovereign AI deployments.
Aleph Alpha says the English-German MoE model has 78.1B total parameters, activates 3.46B per token, and can be served with up to 1,048,576 tokens of context. The company trained it on nearly 24T tokens using B200 GPUs, with architecture choices meant to limit inference cost. Its benchmark claims are vendor-run, including strong scores on math, coding, grounding, and bilingual tests. Grounding is a stated focus, with abstention training meant to make the model withhold answers when evidence is missing. TestingCatalog's note
Aleph Alpha says the English-German MoE model has 78.1B total parameters, activates 3.46B per token, and can be served with up to 1,048,576 tokens of context. The company trained it on nearly 24T tokens using B200 GPUs, with architecture choices meant to limit inference cost. Its benchmark claims are vendor-run, including strong scores on math, coding, grounding, and bilingual tests. Grounding is a stated focus, with abstention training meant to make the model withhold answers when evidence is missing. TestingCatalog's note
score 7