Anything not like your other LLM Stack – Part 1

Hi Guys,

Today, I’ll be presenting a brand new LLM that you can configure locally & based on your computation, I/O & other factors. We’re going to understand the open-source project named -> Colibri.

We’ll be presenting a series of simple & complex process, which will be partially optimized based on various scenarious.

Colibrì is an open-source inference engine for running very large Mixture-of-Experts (MoE) AI models locally. Its central idea is unusual: instead of requiring hundreds of gigabytes—or even terabytes—of model weights to fit entirely in RAM or GPU VRAM, it treats VRAM + RAM + NVMe storage as one memory hierarchy.

The router determines which experts are needed for each token, and Colibrì loads or caches those experts as required. In other words, it effectively performs just-in-time loading of model weights.


Let us understand the overall capability in a simpler form –

CapabilityWhat Colibrì provides
Local large-model inferenceRuns very large MoE models locally rather than requiring a hosted API.
Models larger than RAM/VRAMStreams routed experts from SSD/NVMe instead of loading the full model into memory.
CPU-only operationA GPU is not mandatory. CPU-only operation is supported.
GPU accelerationSupports heterogeneous acceleration paths including CUDA, Metal, Vulkan, and AMD/HIP in documented configurations.
Memory multitieringModel weights can be distributed across VRAM → RAM → NVMe according to available hardware.
Large open modelsCurrently documents support for nine MoE model families, including GLM, Kimi, DeepSeek, Qwen, Inkling, and OLMoE.
Local Chat interfacecoli chat provides direct interactive inference.
Web UIcoli web provides a browser-based interface with performance information and model visualization.
OpenAI-compatible APIcoli serve exposes endpoints such as /v1/chat/completions and /v1/models.
Anthropic-compatible APIAlso provides /v1/messages, allowing some Anthropic-compatible software to communicate with Colibrì.
Tool/function callingAvailable for certain supported engines, including GLM, DeepSeek V4, and Kimi K3.
Coding-tool integrationCan act as a local backend for compatible coding applications that accept an OpenAI-compatible endpoint.
Persistent KV cacheConversation state can be persisted and reused rather than completely prefilling the model again.
Multi-machine experimentsIncludes an experimental/local cluster mechanism in which other Macs can execute routed expert FFNs.
Research/observabilityProvides expert routing telemetry, profiling, Brain visualization, Atlas visualization, benchmarking, and experimentation facilities.
Open sourceEngine is released under Apache 2.0.

For more information, please refer to the following link.

Let us understand how this works for the larger models –

For the project’s reference GLM-5.2 model, the approximately 430 GB quantized model does not need 430 GB of RAM. The documentation describes keeping roughly 9.9 GB of dense components resident while streaming routed experts from storage as needed.

Please go throught he following link for more details.

Let us calculate it –

Model → must largely fit RAM/VRAM → inference
NVMe ↔ RAM ↔ VRAM → dynamically selected experts → inference

Colibrì calls this a multitier hierarchy. Fast memory affects performance, but the design aims to avoid changing the model merely because less VRAM/RAM is available. For more information, let us understand it from the official link.

That architectural concept is probably the project’s most significant contribution.

Model familyApprox. total parametersApprox. storage/RAM characteristics documented by Colibrì
GLM-5.2 / GLM-5.3744B~372 GB disk; 16 GB minimum RAM
GLM-5.3-Flash321B~195 GB converted; ~25 GB RAM
Inkling975B~469 GB; ~25 GB RAM with int4 dense container
Kimi K32.8T~1.6 TB disk; 32 GB+ RAM
DeepSeek V4 Flash284B~167 GB; 16 GB minimum
DeepSeek V4.1 Flash552Bseparate current engine implementation
Qwen3.8-Flash-Next125B + 51B n-gram~185.5 GB; 16 GB minimum
Qwen3.6-35B-A3B35B / ~3B active~20 GB; ~24 GB RAM
OLMoE7B~7 GB; ~8 GB RAM

The current repository specifically describes Kimi K3 at 2.8 trillion parameters, which illustrates how far the storage-streaming approach is intended to scale. One important distinction: Colibrì does not create these models. It is the inference engine that runs compatible open-weight models.

This is another major difference from frameworks primarily designed around GPU residency. The official quick-start says the baseline requirements for GLM-5.2 are approximately:

~16 GB RAM minimum, ~380 GB free storage, with fast NVMe recommended.

A GPU is optional. Refer to this link.

A GPU can significantly accelerate inference, but Colibrì’s architecture intentionally allows CPU + RAM + SSD operation. The repository currently contains backends for CUDA, Metal, Vulkan, and AMD/HIP-related configurations. Please refer the following link.

This makes Colibrì particularly interesting for heterogeneous hardware rather than only expensive data-center GPU configurations.

So, you don’t need the GPU to run any model.


This could be one of its most useful practical capabilities.

Run the following command:

coli serve

Creates an OpenAI-compatible HTTP API. The official documentation currently lists endpoints, including:

GET /v1/models

GET /v1/models/{model}

POST /v1/chat/completions

POST /v1/completions

It supports streaming responses, token usage, temperature, top-p, maximum-token parameters, and stop sequences.

So, you can use the application either from the postman or using API-based call or from using a curl statement while invoking the following entries –

http://localhost:8000/v1

That makes Colibrì much more than a command-line model runner.

Interestingly, Colibrì now exposes:

/v1/messages

for software built around the Anthropic Messages API.

The documentation specifically discusses using this interface with compatible clients, including Claude Code-style integrations. For more information, please refer the following link.

Therefore, the architecture can look approximately like:

This opens some interesting possibilities for private local AI applications, coding assistants, RAG systems, research environments, and experimental agents.


Some of the models support tool/function calling. This is particularly relevent if you are thinking of agentic systems. So, Colibri potentially provides the LLM reasoing/tool-call layer of a local agent architecture.

Colibrì itself isn’t a complete agent orchestration framework like AutoGen or similar systems; it provides the model inference/API layer on which an agent framework could operate.

Colibri also offered web dashboard accessible through:

coli web

The project describes three particularly interesting views –

Dashboard provides token-generation and performance metrics, including time breakdowns and VRAM/RAM/disk utilization.

Brain visualizes the model’s experts and shows which experts activate during generation.

Atlas visualizes measured expert specialization and routing affinities as a three-dimensional representation.

That makes Colibrì interesting not only for inference but also for studying how MoE models route information internally.

For more information, you can refer the following link.


Let us continue this topic in our next series.

Until then, Happy Avenging! 🙂

One thought on “Anything not like your other LLM Stack – Part 1”

Leave a Reply