Hi Guys,
Today, I’ll be presenting a brand new LLM that you can configure locally & based on your computation, I/O & other factors. We’re going to understand the open-source project named -> Colibri.
We’ll be presenting a series of simple & complex process, which will be partially optimized based on various scenarious.
Colibrì is an open-source inference engine for running very large Mixture-of-Experts (MoE) AI models locally. Its central idea is unusual: instead of requiring hundreds of gigabytes—or even terabytes—of model weights to fit entirely in RAM or GPU VRAM, it treats VRAM + RAM + NVMe storage as one memory hierarchy.
The router determines which experts are needed for each token, and Colibrì loads or caches those experts as required. In other words, it effectively performs just-in-time loading of model weights.
Let us understand the overall capability in a simpler form –
| Capability | What Colibrì provides |
|---|---|
| Local large-model inference | Runs very large MoE models locally rather than requiring a hosted API. |
| Models larger than RAM/VRAM | Streams routed experts from SSD/NVMe instead of loading the full model into memory. |
| CPU-only operation | A GPU is not mandatory. CPU-only operation is supported. |
| GPU acceleration | Supports heterogeneous acceleration paths including CUDA, Metal, Vulkan, and AMD/HIP in documented configurations. |
| Memory multitiering | Model weights can be distributed across VRAM → RAM → NVMe according to available hardware. |
| Large open models | Currently documents support for nine MoE model families, including GLM, Kimi, DeepSeek, Qwen, Inkling, and OLMoE. |
| Local Chat interface | coli chat provides direct interactive inference. |
| Web UI | coli web provides a browser-based interface with performance information and model visualization. |
| OpenAI-compatible API | coli serve exposes endpoints such as /v1/chat/completions and /v1/models. |
| Anthropic-compatible API | Also provides /v1/messages, allowing some Anthropic-compatible software to communicate with Colibrì. |
| Tool/function calling | Available for certain supported engines, including GLM, DeepSeek V4, and Kimi K3. |
| Coding-tool integration | Can act as a local backend for compatible coding applications that accept an OpenAI-compatible endpoint. |
| Persistent KV cache | Conversation state can be persisted and reused rather than completely prefilling the model again. |
| Multi-machine experiments | Includes an experimental/local cluster mechanism in which other Macs can execute routed expert FFNs. |
| Research/observability | Provides expert routing telemetry, profiling, Brain visualization, Atlas visualization, benchmarking, and experimentation facilities. |
| Open source | Engine is released under Apache 2.0. |
For more information, please refer to the following link.
Why large models works here?
Let us understand how this works for the larger models –
For the project’s reference GLM-5.2 model, the approximately 430 GB quantized model does not need 430 GB of RAM. The documentation describes keeping roughly 9.9 GB of dense components resident while streaming routed experts from storage as needed.
Please go throught he following link for more details.
Let us calculate it –
Traditional local inference:
Model → must largely fit RAM/VRAM → inference
Colibrì:
NVMe ↔ RAM ↔ VRAM → dynamically selected experts → inference
Colibrì calls this a multitier hierarchy. Fast memory affects performance, but the design aims to avoid changing the model merely because less VRAM/RAM is available. For more information, let us understand it from the official link.
That architectural concept is probably the project’s most significant contribution.
Supported Models:
| Model family | Approx. total parameters | Approx. storage/RAM characteristics documented by Colibrì |
|---|---|---|
| GLM-5.2 / GLM-5.3 | 744B | ~372 GB disk; 16 GB minimum RAM |
| GLM-5.3-Flash | 321B | ~195 GB converted; ~25 GB RAM |
| Inkling | 975B | ~469 GB; ~25 GB RAM with int4 dense container |
| Kimi K3 | 2.8T | ~1.6 TB disk; 32 GB+ RAM |
| DeepSeek V4 Flash | 284B | ~167 GB; 16 GB minimum |
| DeepSeek V4.1 Flash | 552B | separate current engine implementation |
| Qwen3.8-Flash-Next | 125B + 51B n-gram | ~185.5 GB; 16 GB minimum |
| Qwen3.6-35B-A3B | 35B / ~3B active | ~20 GB; ~24 GB RAM |
| OLMoE | 7B | ~7 GB; ~8 GB RAM |
The current repository specifically describes Kimi K3 at 2.8 trillion parameters, which illustrates how far the storage-streaming approach is intended to scale. One important distinction: Colibrì does not create these models. It is the inference engine that runs compatible open-weight models.
Do you need GPU to run this?
This is another major difference from frameworks primarily designed around GPU residency. The official quick-start says the baseline requirements for GLM-5.2 are approximately:
~16 GB RAM minimum, ~380 GB free storage, with fast NVMe recommended.
A GPU is optional. Refer to this link.
A GPU can significantly accelerate inference, but Colibrì’s architecture intentionally allows CPU + RAM + SSD operation. The repository currently contains backends for CUDA, Metal, Vulkan, and AMD/HIP-related configurations. Please refer the following link.
This makes Colibrì particularly interesting for heterogeneous hardware rather than only expensive data-center GPU configurations.
So, you don’t need the GPU to run any model.
Use it like a OpenAI Server:
This could be one of its most useful practical capabilities.
Run the following command:
coli serveCreates an OpenAI-compatible HTTP API. The official documentation currently lists endpoints, including:
GET /v1/models
GET /v1/models/{model}
POST /v1/chat/completions
POST /v1/completionsIt supports streaming responses, token usage, temperature, top-p, maximum-token parameters, and stop sequences.
So, you can use the application either from the postman or using API-based call or from using a curl statement while invoking the following entries –
http://localhost:8000/v1That makes Colibrì much more than a command-line model runner.
Use it like a Anthropic Server:
Interestingly, Colibrì now exposes:
/v1/messagesfor software built around the Anthropic Messages API.
The documentation specifically discusses using this interface with compatible clients, including Claude Code-style integrations. For more information, please refer the following link.
Therefore, the architecture can look approximately like:

This opens some interesting possibilities for private local AI applications, coding assistants, RAG systems, research environments, and experimental agents.
Some of the models support tool/function calling. This is particularly relevent if you are thinking of agentic systems. So, Colibri potentially provides the LLM reasoing/tool-call layer of a local agent architecture.

Colibrì itself isn’t a complete agent orchestration framework like AutoGen or similar systems; it provides the model inference/API layer on which an agent framework could operate.
Colibri also offered web dashboard accessible through:
coli webThe project describes three particularly interesting views –
Dashboard provides token-generation and performance metrics, including time breakdowns and VRAM/RAM/disk utilization.
Brain visualizes the model’s experts and shows which experts activate during generation.
Atlas visualizes measured expert specialization and routing affinities as a three-dimensional representation.
That makes Colibrì interesting not only for inference but also for studying how MoE models route information internally.
For more information, you can refer the following link.
Let us continue this topic in our next series.
Until then, Happy Avenging! 🙂
Note: All the data & scenarios posted here are representative of data & scenarios available on the internet for educational purposes only. There is always room for improvement in this kind of model & the solution associated with it. This article is for educational purposes only. The techniques described should only be used for authorized security testing and research. Unauthorized access to computer systems is illegal and unethical & not encouraged.
One thought on “Anything not like your other LLM Stack – Part 1”