🆕 Haystack 3.3 is here! Faster SentenceWindowRetriever without code changes, bug fixes and a security update
Maintained by deepset

Integration: Kolibri

Use Aleph Alpha Kolibri open-weight German–English models with Haystack for generation, reasoning, and agents

Authors
deepset

Table of Contents

Overview

Kolibri is Aleph Alpha’s open-weight model for German and English. Serve it with vLLM, then use it in Haystack for chat, reasoning, and agents.

Key capabilities:

Capability Details
Generative Chat completion for drafting, Q&A, coding, and long-document work
German & English Native bilingual model; German is a first-class language (≈24% of pre-training mix)
Reasoning Controllable thinking via reasoning_effort (none / low / medium / high)
Agentic Hermes-style tool calling with vLLM’s kolibri1 parser
Long context Native 262,144 tokens; validated up to 1,048,576 (recommend ≤256k for complex tasks)

Setup

Kolibri needs Aleph Alpha’s vLLM plugin (aleph-alpha-inference). Stock vLLM alone cannot load Kolibri1ForCausalLM or the kolibri1 parsers.

Option A — pip (pins a supported vLLM, currently 0.29):

pip install 'aleph-alpha-inference>=1'

Option B — container:

docker pull ghcr.io/aleph-alpha/aleph-alpha-inference

Serve with reasoning and tool calling enabled:

vllm serve Aleph-Alpha/Kolibri-1 \
  --kv-cache-dtype fp8 \
  --reasoning-parser kolibri1 \
  --tool-call-parser kolibri1 \
  --enable-auto-tool-choice

For BF16 weights, serve Aleph-Alpha/Kolibri-1-BF16 and omit --kv-cache-dtype fp8. For contexts beyond 262,144 tokens, add --max-model-len 1048576 --hf-overrides '{"max_position_embeddings": 1048576}'.

Recommended sampling defaults from the model card: temperature=1.0, top_p=0.97, top_k=128.

Install the Haystack client for the OpenAI-compatible vLLM server:

pip install vllm-haystack

See also the vLLM integration.

Usage

Once Kolibri is serving (default http://localhost:8000/v1), use VLLMChatGenerator for generation, reasoning, and Haystack Agent agent loops.

Generative chat (German)

Kolibri is specialized for German. This example asks for a short explanation entirely in German:

from haystack.dataclasses import ChatMessage
from haystack_integrations.components.generators.vllm import VLLMChatGenerator

chat = VLLMChatGenerator(
    model="Aleph-Alpha/Kolibri-1",
    api_base_url="http://localhost:8000/v1",
    generation_kwargs={
        "temperature": 1.0,
        "top_p": 0.97,
        "extra_body": {
            "top_k": 128,
            "chat_template_kwargs": {
                "enable_thinking": False,
                "reasoning_effort": "none",
            },
        },
    },
)

result = chat.run(
    messages=[
        ChatMessage.from_user(
            "Erkläre kurz, was ein Mixture-of-Experts-Modell ist und warum es für deutsche Verwaltungstexte hilfreich sein kann."
        )
    ]
)
print(result["replies"][0].text)

Reasoning mode

Thinking is on by default when the server is started with --reasoning-parser kolibri1. Control effort per request with chat_template_kwargs. When thinking is enabled, Haystack exposes the reasoning trace on the reply’s .reasoning field:

from haystack.dataclasses import ChatMessage
from haystack_integrations.components.generators.vllm import VLLMChatGenerator

chat = VLLMChatGenerator(
    model="Aleph-Alpha/Kolibri-1",
    api_base_url="http://localhost:8000/v1",
    generation_kwargs={
        "temperature": 1.0,
        "top_p": 0.97,
        "extra_body": {
            "top_k": 128,
            "chat_template_kwargs": {
                "enable_thinking": True,
                "reasoning_effort": "high",  # none | low | medium | high
            },
        },
    },
)

result = chat.run(
    messages=[
        ChatMessage.from_user(
            "Ein Amt hat 3 Anträge: A braucht 2 Tage, B 5 Tage, C 1 Tag. "
            "In welcher Reihenfolge minimiert man die mittlere Wartezeit? Begründe."
        )
    ]
)
reply = result["replies"][0]
print(reply.reasoning)  # thinking trace (when enabled)
print(reply.text)       # final answer

Set reasoning_effort to "none" or enable_thinking to False for a direct reply without a thinking block.

Agentic tool calling

Serve with --tool-call-parser kolibri1 --enable-auto-tool-choice, then plug Kolibri into a Haystack Agent. Tool calling can be combined with reasoning. This agent looks up a mock product price in German:

from typing import Annotated

from haystack.components.agents import Agent
from haystack.dataclasses import ChatMessage
from haystack.tools import tool
from haystack_integrations.components.generators.vllm import VLLMChatGenerator


@tool
def preis_abrufen(produkt: Annotated[str, "Produktname für die Preissuche"]) -> str:
    """Sucht einen Beispielpreis für ein Produkt."""
    preise = {"laptop": "999 €", "tastatur": "99 €", "monitor": "349 €"}
    return preise.get(produkt.lower(), "Unbekanntes Produkt")


agent = Agent(
    chat_generator=VLLMChatGenerator(
        model="Aleph-Alpha/Kolibri-1",
        api_base_url="http://localhost:8000/v1",
        generation_kwargs={
            "temperature": 1.0,
            "top_p": 0.97,
            "extra_body": {
                "top_k": 128,
                "chat_template_kwargs": {
                    "enable_thinking": True,
                    "reasoning_effort": "medium",
                },
            },
        },
    ),
    tools=[preis_abrufen],
    system_prompt=(
        "Du hilfst Nutzern auf Deutsch und rufst die bereitgestellten Tools auf, wenn sie relevant sind. Antworte klar und knapp."
    ),
)

result = agent.run(messages=[ChatMessage.from_user("Was kostet ein Laptop?")])
print(result["last_message"].text)

For English-only assistants, the same pattern works — keep tools and prompts in English. For RAG pipelines, use Kolibri as the generator after retrieval the same way you would any other chat model behind vLLM.