
This Repo Runs a 26B Model in 2GB of RAM, 4k Stars
Bitwise AI · 6:21
Turbo Fieldfare runs Google's Gemma 4 26B (MoE, 3.9B active params per token) on an 8GB M2 MacBook Air at ~6 tok/s by keeping only 1.35GB of shared weights in RAM and loading routed experts from SSD on demand — explicit parallel reads replaced mmap for a 7x speedup. Caveats: OS file cache does heavy lifting behind the 2GB footprint, 4-bit quantized weights with no published quality comparison, and MLX is 2.5x faster if you have 14GB free.
Alösha's take: Finally someone shipped what Apple Research proposed in 2023 — SSD as the new VRAM. The engineering is clean and the honest benchmarks (with asterisks) make this worth studying.



