Anything not like your other LLM Stack – Part 2

In the part – 1, we’ve discussed the details on the basic concept of this enw LLM frame work & how it enables the similar capability without the need for accelerated GPUs or specific hardwares. In this post, we’ll further discuss this & explore more details on top of that.


The engine records routing activity and can keep frequently used experts in faster memory. It combines mechanisms such as:

NVMe → RAM caching → pinned hot experts → optional GPU residency → prefetching

The project describes this as analogous to a JIT compiler for weights: instead of compiling frequently executed code, it learns which model experts are frequently requested and moves those experts into faster storage tiers. 

This is why repeated workloads can behave differently from completely cold inference.


Colibrì can also distribute expert reads across two storage devices.

Because disk bandwidth is frequently the limiting factor, the engine supports a second model mirror and routes reads across the drives according to measured or configured bandwidth.

And, we’ll be also exploring this appraoch to reduce the load in our MacBook Pro.

So something conceptually like:

                     ┌── NVMe SSD #1
Model expert request ┤
└── NVMe SSD #2
↓
RAM
↓
GPU

This can increase the effective streaming bandwidth.

That is a fairly distinctive capability compared with conventional local LLM runtimes.

The API also supports persistent KV contexts.

Colibrì can save KV state to .coli_kv files and reuse the common prefix when a conversation continues. For more information please refer the following link.

That can matter enormously for huge models because repeatedly processing a long conversation history could otherwise be extremely expensive.

Multiple isolated KV slots are also supported, currently up to 16 according to the API documentation. Refer the below link to know more on this.

Colibrì demonstrates that enormous models can run on modest hardware.

That does not mean they run at cloud-API speeds. The project’s own benchmark documentation is unusually transparent about this.

For the 744B GLM model, documented examples include roughly:

  • 25 GB development machine, cold: 0.05–0.1 tokens/s
  • 128 GB CPU-only system, warm: ~1.8 tokens/s
  • RTX 5070 Ti system: ~1.07 tokens/s in one documented configuration
  • Large six-RTX-5090 configuration: several tokens per second with experts resident rather than coming from disk.

In Colibri, the coding agent might send 10,000–20,000 tokens before the user’s first message, and disk-streamed prefill could therefore take an extremely long time.

Colibrì’s innovation is primarily accessibility and heterogeneous resource utilization, not magically eliminating the physics of memory bandwidth.

The key contribution is not simply “running an LLM locally.” Tools such as Ollama already make that common.

Colibrì is exploring a different question:

What if model size no longer had to be constrained by available RAM or VRAM?

The repository treats storage, memory, caching, routing patterns, speculative decoding, quantization, NUMA, CPU/GPU overlap, and expert placement as parts of the inference system itself.

That could make it valuable for four broad areas:

  • Local/private inference, particularly when sending data to an external inference provider is undesirable.
  • AI systems research, especially research into MoE routing, caching, memory hierarchy, quantization, and heterogeneous computing.
  • Low-cost experimentation with frontier-scale open models, where owning several data-center GPUs would otherwise be required.
  • Building local AI services, because its OpenAI/Anthropic-compatible APIs can sit underneath applications, RAG systems, coding tools, or agent frameworks.

One qualification is important: local execution can improve control over where prompts and model execution occur, but “local” by itself should not be interpreted as an automatic security or privacy guarantee. Network exposure, APIs, applications connected to the server, file access, and logging still need to be configured appropriately. Colibrì defaults the server to localhost and recommends setting an API key before exposing it beyond the machine.


Now, let us understand how we want to proceed with our architecture implementation locally even further reducing the need for the SSD memories of my Macbook Pro.

Colibrì supports a useful feature called a partial model mirror. You can retain the complete model on your external drive like mine (SD_BLACK) and copy a selected portion of frequently used model shards onto your Mac’s much faster internal SSD. Colibrì then reads from both locations. Its official documentation provides commands to plan, stage and verify these partial mirrors. I tried both the options, partially, stored in Macbook & partially stored in my high speed SSD USB. And, alo test by placing the complete model in my MacBook Pro to understand the implications of the performance for storing them in different places.

Now, what my external drive will tell us about the performance –

I’ve used a faster SSD USB from SanDisk to get the better response, while storing the Colibri library over there. That way I can get the almost similar response from the external SSD USB like my internal Macbook Pro’s SSD.

SanDisk officially specifies the Extreme Portable SSD at up to 1,050 MB/s sequential read and 1,000 MB/s sequential write, using USB 3.2 Gen 2. For more on this device, you can refer the following link.

Let us understand, how the cache works for us –

From thje above image, you can see that by default, the main models stays in the faster SanDisk SSD USB. However, 100 GB of the mirror models are kept inside the main laptop’s SSD for even better performance. So, it is a mid-level tradeoff againsts the lapotp’s memory utilization Vs performance.

Using the following way, you can configure the model endpoint ->

export COLI_MODEL="/Volumes/SD_BLACK/AI/models/glm52_i4"
export MIRROR="$HOME/ColibriMirror/glm52_i4"

df -h "$HOME"
mkdir -p "$MIRROR"

cd /Volumes/SD_BLACK/AI/colibri/c

./coli mirror plan \
  --model "$COLI_MODEL" \
  --mirror "$MIRROR" \
  --budget-gib 100 \
  --reserve-gib 60

Review the plan first. Provided you have sufficient internal SSD space, stage and verify the selected shards:

./coli mirror stage \
  --model "$COLI_MODEL" \
  --mirror "$MIRROR" \
  --budget-gib 100 \
  --reserve-gib 60

./coli mirror verify \
  --model "$COLI_MODEL" \
  --mirror "$MIRROR"

These commands are supported by the current repository; an older local Colibrì executable may need updating if it does not recognize coli mirror. Staging preserves the original external model and verifies copied shards. For more information on this, please refer to the following link.

After successful verification, start the server with the mirror enabled:

COLI_MODEL_MIRROR="$MIRROR" \
COLI_METAL=1 \
CAP_RAISE=0 \
MTP=0 \
./coli serve \
  --model "$COLI_MODEL" \
  --ram 88 \
  --cap 8 \
  --host 127.0.0.1 \
  --port 8000 \
  --model-id glm-5.2-colibri

Now, in the next thread we’ll understand the entire step-by-step setup of this initiative & explian one more complex combinations to get even better resolution with custom twists.

Till then, Happy Avenging! 😀

Anything not like your other LLM Stack – Part 1

Hi Guys,

Today, I’ll be presenting a brand new LLM that you can configure locally & based on your computation, I/O & other factors. We’re going to understand the open-source project named -> Colibri.

We’ll be presenting a series of simple & complex process, which will be partially optimized based on various scenarious.

Colibrì is an open-source inference engine for running very large Mixture-of-Experts (MoE) AI models locally. Its central idea is unusual: instead of requiring hundreds of gigabytes—or even terabytes—of model weights to fit entirely in RAM or GPU VRAM, it treats VRAM + RAM + NVMe storage as one memory hierarchy.

The router determines which experts are needed for each token, and Colibrì loads or caches those experts as required. In other words, it effectively performs just-in-time loading of model weights.


Let us understand the overall capability in a simpler form –

CapabilityWhat Colibrì provides
Local large-model inferenceRuns very large MoE models locally rather than requiring a hosted API.
Models larger than RAM/VRAMStreams routed experts from SSD/NVMe instead of loading the full model into memory.
CPU-only operationA GPU is not mandatory. CPU-only operation is supported.
GPU accelerationSupports heterogeneous acceleration paths including CUDA, Metal, Vulkan, and AMD/HIP in documented configurations.
Memory multitieringModel weights can be distributed across VRAM → RAM → NVMe according to available hardware.
Large open modelsCurrently documents support for nine MoE model families, including GLM, Kimi, DeepSeek, Qwen, Inkling, and OLMoE.
Local Chat interfacecoli chat provides direct interactive inference.
Web UIcoli web provides a browser-based interface with performance information and model visualization.
OpenAI-compatible APIcoli serve exposes endpoints such as /v1/chat/completions and /v1/models.
Anthropic-compatible APIAlso provides /v1/messages, allowing some Anthropic-compatible software to communicate with Colibrì.
Tool/function callingAvailable for certain supported engines, including GLM, DeepSeek V4, and Kimi K3.
Coding-tool integrationCan act as a local backend for compatible coding applications that accept an OpenAI-compatible endpoint.
Persistent KV cacheConversation state can be persisted and reused rather than completely prefilling the model again.
Multi-machine experimentsIncludes an experimental/local cluster mechanism in which other Macs can execute routed expert FFNs.
Research/observabilityProvides expert routing telemetry, profiling, Brain visualization, Atlas visualization, benchmarking, and experimentation facilities.
Open sourceEngine is released under Apache 2.0.

For more information, please refer to the following link.

Let us understand how this works for the larger models –

For the project’s reference GLM-5.2 model, the approximately 430 GB quantized model does not need 430 GB of RAM. The documentation describes keeping roughly 9.9 GB of dense components resident while streaming routed experts from storage as needed.

Please go throught he following link for more details.

Let us calculate it –

Model → must largely fit RAM/VRAM → inference
NVMe ↔ RAM ↔ VRAM → dynamically selected experts → inference

Colibrì calls this a multitier hierarchy. Fast memory affects performance, but the design aims to avoid changing the model merely because less VRAM/RAM is available. For more information, let us understand it from the official link.

That architectural concept is probably the project’s most significant contribution.

Model familyApprox. total parametersApprox. storage/RAM characteristics documented by Colibrì
GLM-5.2 / GLM-5.3744B~372 GB disk; 16 GB minimum RAM
GLM-5.3-Flash321B~195 GB converted; ~25 GB RAM
Inkling975B~469 GB; ~25 GB RAM with int4 dense container
Kimi K32.8T~1.6 TB disk; 32 GB+ RAM
DeepSeek V4 Flash284B~167 GB; 16 GB minimum
DeepSeek V4.1 Flash552Bseparate current engine implementation
Qwen3.8-Flash-Next125B + 51B n-gram~185.5 GB; 16 GB minimum
Qwen3.6-35B-A3B35B / ~3B active~20 GB; ~24 GB RAM
OLMoE7B~7 GB; ~8 GB RAM

The current repository specifically describes Kimi K3 at 2.8 trillion parameters, which illustrates how far the storage-streaming approach is intended to scale. One important distinction: Colibrì does not create these models. It is the inference engine that runs compatible open-weight models.

This is another major difference from frameworks primarily designed around GPU residency. The official quick-start says the baseline requirements for GLM-5.2 are approximately:

~16 GB RAM minimum, ~380 GB free storage, with fast NVMe recommended.

A GPU is optional. Refer to this link.

A GPU can significantly accelerate inference, but Colibrì’s architecture intentionally allows CPU + RAM + SSD operation. The repository currently contains backends for CUDA, Metal, Vulkan, and AMD/HIP-related configurations. Please refer the following link.

This makes Colibrì particularly interesting for heterogeneous hardware rather than only expensive data-center GPU configurations.

So, you don’t need the GPU to run any model.


This could be one of its most useful practical capabilities.

Run the following command:

coli serve

Creates an OpenAI-compatible HTTP API. The official documentation currently lists endpoints, including:

GET /v1/models

GET /v1/models/{model}

POST /v1/chat/completions

POST /v1/completions

It supports streaming responses, token usage, temperature, top-p, maximum-token parameters, and stop sequences.

So, you can use the application either from the postman or using API-based call or from using a curl statement while invoking the following entries –

http://localhost:8000/v1

That makes Colibrì much more than a command-line model runner.

Interestingly, Colibrì now exposes:

/v1/messages

for software built around the Anthropic Messages API.

The documentation specifically discusses using this interface with compatible clients, including Claude Code-style integrations. For more information, please refer the following link.

Therefore, the architecture can look approximately like:

This opens some interesting possibilities for private local AI applications, coding assistants, RAG systems, research environments, and experimental agents.


Some of the models support tool/function calling. This is particularly relevent if you are thinking of agentic systems. So, Colibri potentially provides the LLM reasoing/tool-call layer of a local agent architecture.

Colibrì itself isn’t a complete agent orchestration framework like AutoGen or similar systems; it provides the model inference/API layer on which an agent framework could operate.

Colibri also offered web dashboard accessible through:

coli web

The project describes three particularly interesting views –

Dashboard provides token-generation and performance metrics, including time breakdowns and VRAM/RAM/disk utilization.

Brain visualizes the model’s experts and shows which experts activate during generation.

Atlas visualizes measured expert specialization and routing affinities as a three-dimensional representation.

That makes Colibrì interesting not only for inference but also for studying how MoE models route information internally.

For more information, you can refer the following link.


Let us continue this topic in our next series.

Until then, Happy Avenging! 🙂