Anything not like your other LLM Stack – Part 2

In the part – 1, we’ve discussed the details on the basic concept of this enw LLM frame work & how it enables the similar capability without the need for accelerated GPUs or specific hardwares. In this post, we’ll further discuss this & explore more details on top of that.


The engine records routing activity and can keep frequently used experts in faster memory. It combines mechanisms such as:

NVMe → RAM caching → pinned hot experts → optional GPU residency → prefetching

The project describes this as analogous to a JIT compiler for weights: instead of compiling frequently executed code, it learns which model experts are frequently requested and moves those experts into faster storage tiers. 

This is why repeated workloads can behave differently from completely cold inference.


Colibrì can also distribute expert reads across two storage devices.

Because disk bandwidth is frequently the limiting factor, the engine supports a second model mirror and routes reads across the drives according to measured or configured bandwidth.

And, we’ll be also exploring this appraoch to reduce the load in our MacBook Pro.

So something conceptually like:

                     ┌── NVMe SSD #1
Model expert request ┤
└── NVMe SSD #2
↓
RAM
↓
GPU

This can increase the effective streaming bandwidth.

That is a fairly distinctive capability compared with conventional local LLM runtimes.

The API also supports persistent KV contexts.

Colibrì can save KV state to .coli_kv files and reuse the common prefix when a conversation continues. For more information please refer the following link.

That can matter enormously for huge models because repeatedly processing a long conversation history could otherwise be extremely expensive.

Multiple isolated KV slots are also supported, currently up to 16 according to the API documentation. Refer the below link to know more on this.

Colibrì demonstrates that enormous models can run on modest hardware.

That does not mean they run at cloud-API speeds. The project’s own benchmark documentation is unusually transparent about this.

For the 744B GLM model, documented examples include roughly:

  • 25 GB development machine, cold: 0.05–0.1 tokens/s
  • 128 GB CPU-only system, warm: ~1.8 tokens/s
  • RTX 5070 Ti system: ~1.07 tokens/s in one documented configuration
  • Large six-RTX-5090 configuration: several tokens per second with experts resident rather than coming from disk.

In Colibri, the coding agent might send 10,000–20,000 tokens before the user’s first message, and disk-streamed prefill could therefore take an extremely long time.

Colibrì’s innovation is primarily accessibility and heterogeneous resource utilization, not magically eliminating the physics of memory bandwidth.

The key contribution is not simply “running an LLM locally.” Tools such as Ollama already make that common.

Colibrì is exploring a different question:

What if model size no longer had to be constrained by available RAM or VRAM?

The repository treats storage, memory, caching, routing patterns, speculative decoding, quantization, NUMA, CPU/GPU overlap, and expert placement as parts of the inference system itself.

That could make it valuable for four broad areas:

  • Local/private inference, particularly when sending data to an external inference provider is undesirable.
  • AI systems research, especially research into MoE routing, caching, memory hierarchy, quantization, and heterogeneous computing.
  • Low-cost experimentation with frontier-scale open models, where owning several data-center GPUs would otherwise be required.
  • Building local AI services, because its OpenAI/Anthropic-compatible APIs can sit underneath applications, RAG systems, coding tools, or agent frameworks.

One qualification is important: local execution can improve control over where prompts and model execution occur, but “local” by itself should not be interpreted as an automatic security or privacy guarantee. Network exposure, APIs, applications connected to the server, file access, and logging still need to be configured appropriately. Colibrì defaults the server to localhost and recommends setting an API key before exposing it beyond the machine.


Now, let us understand how we want to proceed with our architecture implementation locally even further reducing the need for the SSD memories of my Macbook Pro.

Colibrì supports a useful feature called a partial model mirror. You can retain the complete model on your external drive like mine (SD_BLACK) and copy a selected portion of frequently used model shards onto your Mac’s much faster internal SSD. Colibrì then reads from both locations. Its official documentation provides commands to plan, stage and verify these partial mirrors. I tried both the options, partially, stored in Macbook & partially stored in my high speed SSD USB. And, alo test by placing the complete model in my MacBook Pro to understand the implications of the performance for storing them in different places.

Now, what my external drive will tell us about the performance –

I’ve used a faster SSD USB from SanDisk to get the better response, while storing the Colibri library over there. That way I can get the almost similar response from the external SSD USB like my internal Macbook Pro’s SSD.

SanDisk officially specifies the Extreme Portable SSD at up to 1,050 MB/s sequential read and 1,000 MB/s sequential write, using USB 3.2 Gen 2. For more on this device, you can refer the following link.

Let us understand, how the cache works for us –

From thje above image, you can see that by default, the main models stays in the faster SanDisk SSD USB. However, 100 GB of the mirror models are kept inside the main laptop’s SSD for even better performance. So, it is a mid-level tradeoff againsts the lapotp’s memory utilization Vs performance.

Using the following way, you can configure the model endpoint ->

export COLI_MODEL="/Volumes/SD_BLACK/AI/models/glm52_i4"
export MIRROR="$HOME/ColibriMirror/glm52_i4"

df -h "$HOME"
mkdir -p "$MIRROR"

cd /Volumes/SD_BLACK/AI/colibri/c

./coli mirror plan \
  --model "$COLI_MODEL" \
  --mirror "$MIRROR" \
  --budget-gib 100 \
  --reserve-gib 60

Review the plan first. Provided you have sufficient internal SSD space, stage and verify the selected shards:

./coli mirror stage \
  --model "$COLI_MODEL" \
  --mirror "$MIRROR" \
  --budget-gib 100 \
  --reserve-gib 60

./coli mirror verify \
  --model "$COLI_MODEL" \
  --mirror "$MIRROR"

These commands are supported by the current repository; an older local Colibrì executable may need updating if it does not recognize coli mirror. Staging preserves the original external model and verifies copied shards. For more information on this, please refer to the following link.

After successful verification, start the server with the mirror enabled:

COLI_MODEL_MIRROR="$MIRROR" \
COLI_METAL=1 \
CAP_RAISE=0 \
MTP=0 \
./coli serve \
  --model "$COLI_MODEL" \
  --ram 88 \
  --cap 8 \
  --host 127.0.0.1 \
  --port 8000 \
  --model-id glm-5.2-colibri

Now, in the next thread we’ll understand the entire step-by-step setup of this initiative & explian one more complex combinations to get even better resolution with custom twists.

Till then, Happy Avenging! 😀

Hacking the performance of Python Solutions with a custom-built library

Today, I’m very excited to demonstrate an effortless & new way to hack the performance of Python. This post will be a super short & yet crisp presentation of improving the overall performance.

Why not view the demo before going through it?


Demo

Isn’t it exciting? Let’s understand the steps to improve your code.

pip install cython

Cython is a Python-to-C compiler. It can significantly improve performance for specific tasks, especially those with heavy computation and loops. Also, Cython’s syntax is very similar to Python, which makes it easy to learn.

Let’s consider an example where we calculate the sum of squares for a list of numbers. The code without optimization would look like this:

  • perfTest_1.py (First untuned Python class.)
#########################################################
#### Written By: SATYAKI DE                          ####
#### Written On: 31-Jul-2023                         ####
#### Modified On 31-Jul-2023                         ####
####                                                 ####
#### Objective: This is the main calling             ####
#### python script that will invoke the              ####
#### first version of accute computation.            ####
####                                                 ####
#########################################################
from clsConfigClient import clsConfigClient as cf

import time
start = time.time()

n_val = cf.conf['INPUT_VAL']

def compute_sum_of_squares(n):
    return sum([i**2 for i in range(n)])

n = n_val

print(compute_sum_of_squares(n))

print(f"Test - 1: Execution time: {time.time() - start} seconds")

Here, n_val contains the value as – “1000000000”.

Now, let’s optimize it using Cython by installing the abovementioned packages. Then, you will have to create a .pyx file, say “compute.pyx”, with the following code:

cpdef double compute_sum_of_squares(int n):
    return sum([i**2 for i in range(n)])

Now, create a setup.py file to compile it:

###########################################################
#### Written By: SATYAKI DE                            ####
#### Written On: 31-Jul-2023                           ####
#### Modified On 31-Jul-2023                           ####
####                                                   ####
#### Objective: This is the main calling               ####
#### python script that will create the                ####
#### compiled library after executing the compute.pyx. ####
####                                                   ####
###########################################################

from setuptools import setup
from Cython.Build import cythonize

setup(
    ext_modules = cythonize("compute.pyx")
)

Compile it using the command:

python setup.py build_ext --inplace

This will look like the following –

Finally, you can import the function from the compiled “.pyx” file inside the improved code.

  • perfTest_2.py (First untuned Python class.)
#########################################################
#### Written By: SATYAKI DE                          ####
#### Written On: 31-Jul-2023                         ####
#### Modified On 31-Jul-2023                         ####
####                                                 ####
#### Objective: This is the main calling             ####
#### python script that will invoke the              ####
#### optimized & precompiled custom library, which   ####
#### will significantly improve the performance.     ####
####                                                 ####
#########################################################
from clsConfigClient import clsConfigClient as cf
from compute import compute_sum_of_squares

import time
start = time.time()

n_val = cf.conf['INPUT_VAL']

n = n_val

print(compute_sum_of_squares(n))

print(f"Test - 2: Execution time with multiprocessing: {time.time() - start} seconds")

By compiling to C, Cython can speed up loop and function calls, leading to significant speedup for CPU-bound tasks.

Please note that while Cython can dramatically improve performance, it can make the code more complex and harder to debug. Therefore, starting with regular Python and switching to Cython for the performance-critical parts of the code is recommended.


So, finally, we’ve done it. I know that this post is relatively smaller than my earlier post. But, I think, you can get a good hack to improve some of your long-running jobs by applying this trick.

I’ll bring some more exciting topics in the coming days from the Python verse. Please share & subscribe to my post & let me know your feedback.

Till then, Happy Avenging! 🙂