Anything not like your other LLM Stack – Part 3

Welcome back! As you have now understood our objectives for the new initiative that we have explored in the other two posts, this post will show the actual implementation of the framework in a simple step-by-step process. But before that, why don’t you see the final LLM environment and the response from it?

But before that, here are the previous two posts in case you directly land on this page –

Isn’t it exciting?

Great! Now, let us dive into the actual setup & implementation.


Before we proceed, let us understand what the target architecture should look like –

EXTERNAL SSD / NVMe
/Volumes/ColibriSSD/
│
├── AI/
│ ├── colibri/ ← Colibrì source + executable
│ ├── models/
│ │ └── glm52_i4/ ← ~372 GB MODEL
│ │ ├── *.safetensors
│ │ ├── config.json
│ │ ├── tokenizer...
│ │ ├── .coli_usage
│ │ ├── .coli_ssd
│ │ └── .coli_kv
│ │
│ ├── huggingface/ ← HF cache/download metadata
│ ├── config/ ← Colibrì tuning profiles
│ └── venv/ ← optional Python environment
│
│ USB4 / Thunderbolt / USB-C
│ │
│ ▼
│ Your Apple-Silicon Mac
│
│ ┌─────────────────┐
│ │ Unified Memory │
│ │ e.g. 24/48/64GB │
│ └────────┬────────┘
│ │
│ Metal GPU
│ │
│ ▼
│ Colibrì inference
│ │
│ localhost:8000
│ │
│ ┌──────────┼──────────┐
│ ▼ ▼ ▼
│ Python RAG app AI Agent
│ app etc. etc.

Connect the external SSD and run:

ls /Volumes

And the response –

Then define one variable:

export COLI_DRIVE="/Volumes/SD_BLACK"

Verify that macOS sees it:

df -h "$COLI_DRIVE"

And verify available space:

diskutil info "$COLI_DRIVE"

Create a clean directory structure:

mkdir -p "$COLI_DRIVE/AI"
mkdir -p "$COLI_DRIVE/AI/models"
mkdir -p "$COLI_DRIVE/AI/huggingface"
mkdir -p "$COLI_DRIVE/AI/config"

Validate the output –

ls -la "$COLI_DRIVE/AI"

The following components may already exists in your Mac. So, install them based on your availability.

xcode-select --install
brew install libomp git python

Now clone Colibrì directly onto the external SSD:

cd "$COLI_DRIVE/AI"

git clone https://github.com/JustVugg/colibri.git

Your source will therefore be located at:

You can also build it using the following command –

cd "$COLI_DRIVE/AI/colibri/c"

./setup.sh

Because you’re on Apple Silicon, you should use Colibrì’s Metal backend rather than staying CPU-only.

From the above directory, runt he following command:

make colibri METAL=1

You can then confirm that Colibrì exists:

ls -lh colibri

And the output is as follows –

Now define:

export COLI_MODEL="$COLI_DRIVE/AI/models/glm52_i4"

mkdir -p "$COLI_MODEL"

And verify it

echo "$COLI_MODEL"

I would use the project’s currently recommended preconverted GLM-5.2 model instead of downloading the approximately 756 GB FP8 source and converting it unless you specifically want to experiment with quantization.

Colibrì currently recommends:

mastouri/GLM-5.2-colibri-int4-g64-with-int8-mtp

and specifically recommends its group-scaled gs64 configuration.  For more information, you can refer to the following link.

Create an optional Python environment on the external SSD:

python3 -m venv "$COLI_DRIVE/AI/venv"

Activate it:

source "$COLI_DRIVE/AI/venv/bin/activate"

Install the Hugging Face client:

pip install -U huggingface_hub

Now keep Hugging Face’s standard cache on the external drive too:

export HF_HOME="$COLI_DRIVE/AI/huggingface"

HF_HOME controls the root of Hugging Face’s cache; otherwise, it normally defaults to your user cache on the internal drive. 

Now, we are in a position to start downloading the model with the given command –

unset HF_HUB_DISABLE_XET
unset HF_XET_HIGH_PERFORMANCE

export HF_XET_NUM_CONCURRENT_RANGE_GETS=8

HF_HUB_DISABLE_XET=1 \
hf download \
  mastouri/GLM-5.2-colibri-int4-g64-with-int8-mtp \
  --local-dir "$COLI_MODEL" \
  --max-workers 4

And you will see something like the following –

Now, with the following command, you can verify whether all the models downloaded successfully or not –

hf download \
  mastouri/GLM-5.2-colibri-int4-g64-with-int8-mtp \
  --local-dir "$COLI_MODEL" \
  --dry-run

You would see something like this –

As you can see, those models that are already downloaded won’t show the model size beside their name. Otherwise, it will show the name of the model along with its size, as shown below –

Run the following command –

du -sh "$COLI_MODEL"

And you will see something like this –

Then verify the underlying volume:

df -h "$COLI_MODEL"

Now let Colibrì verify the model:

cd "$COLI_DRIVE/AI/colibri/c"

./coli doctor --model "$COLI_MODEL"

Or,

./coli doctor --deep --model "$COLI_MODEL"

doctor --deep validates safetensors headers, tensor layouts, shard completeness, required tensors, and other model-layout details.

Then:

./coli info --model "$COLI_MODEL"

And most importantly:

./coli plan --model "$COLI_MODEL"

coli plan is designed to show the computed disk/RAM/VRAM placement plan. 

This is the command I would use before actually loading the model.

Suppose, purely as an example, you had 64 GB of unified memory.

Instead of telling Colibrì to consume almost everything, you could initially give it:

COLI_METAL=1 \
./coli chat \
  --model "$COLI_MODEL" \
  --ram 48

This means –

~372 GB model
│
└──────────── External SSD
~48 GB max working/cache budget
│
└──────────── Mac unified memory
Compute
│
└──────────── Apple Metal GPU + CPU

The --ram parameter is explicitly Colibrì’s RAM budget for the expert working set; leaving it at 0 lets the software choose an automatic value based on available memory. For more information, you can refer to the following link.

Therefore, you can also start more safely with:

COLI_METAL=1 \
./coli chat \
  --model "$COLI_MODEL"

However, I’ll use some of the more advanced options to get the best performance out of this. “chat” means you will be triggering the chat interface.

This part is particularly useful on a Mac.

On Metal/macOS, Colibrì automatically performs an F_NOCACHE storage test against the model volume and writes the result into:

$COLI_MODEL/.coli_ssd

The file contains the measured throughput and volume identity. Colibrì uses this information when selecting caching behavior.

You can inspect it after the first successful startup:

cat "$COLI_MODEL/.coli_ssd"

would represent a measured storage throughput value used by Colibrì.

This is particularly useful because an external enclosure advertised as “10 Gbps” or “40 Gbps” does not guarantee that the actual SSD delivers that performance under Colibrì’s workload.

For your applications, I recommend running Colibrì as a server rather than launching the model separately for every script.

You need to be in the following path –

cd "$COLI_DRIVE/AI/colibri/c"

Now, set the following variables –

export COLI_DRIVE="/Volumes/SD_BLACK"
export COLI_MODEL="$COLI_DRIVE/AI/models/glm52_i4"
export XDG_CONFIG_HOME="$COLI_DRIVE/AI/config"
export HF_HOME="$COLI_DRIVE/AI/huggingface"

Enable Metal:

export COLI_METAL=1

Then run:

./coli serve \
  --model "$COLI_MODEL" \
  --host 127.0.0.1 \
  --port 8000 \
  --model-id glm-5.2-colibri
  • Note that this will enable the service mode, which can be accessed from any application as an API.

Colibrì’s API documentation defines the OpenAI-compatible base URL as:

http://localhost:8000/v1

and exposes endpoints such as /v1/chat/completions. 

So, the architecture become –

External SSD
│
│ 372 GB model weights
│
▼
Colibrì
│
│ streams active experts
▼
RAM ⇄ Metal GPU
│
▼
localhost:8000/v1
│
┌───┼───────────┬───────────┐
▼ ▼ ▼ ▼
RAG Python AutoGen Your App
App script agent

Now, you can test the application in the following ways –

curl http://127.0.0.1:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model":"glm-5.2-colibri",
    "messages":[
      {
        "role":"user",
        "content":"Explain mixture-of-experts architecture."
      }
    ]
  }'
from openai import OpenAI

client = OpenAI(
    base_url="http://127.0.0.1:8000/v1",
    api_key="local"
)

response = client.chat.completions.create(
    model="glm-5.2-colibri",
    messages=[
        {
            "role": "user",
            "content": "Explain retrieval-augmented generation."
        }
    ]
)

print(response.choices[0].message.content)

Now, let us understand the command that I used –

PROF=1 \
COLI_METAL=1 \
COLI_METAL_PREFILL=1 \
DIRECT=1 \
PIPE=1 \
PIPE_WORKERS=6 \
COLI_NO_OMP_TUNE=1 \
MTP=0 \
./coli chat \
  --model "$COLI_MODEL" \
  --ram 96

And you would see & then ask some questions to get your answer, which should look like this –

Now, let us ask some questions & validate the response –

The performance is quite better. Now, based on your external SSD USB’s speed, you can get closed to your MacBook Pro’s SSD speed.

In the next post, we’ll explore a more advanced combination to explore even better & faster response with the introduction of a custom script. And there will be some upgraded architecture that follows it. However, in this post, we conclude the basic concept of the Colibri framework.


So, we’ve done it.  🙂

I hope you all like this effort & let me know your feedback. I’ll be back with another topic. Until then, Happy Avenging!

Leave a Reply