S-MoE: Seismic Mixture of Experts Engine

S-MoE: Seismic Mixture of Experts Engine

●5 ●24 ●59
calendar_today ago • schedule7 min read

S-MoE: a 235-billion-parameter model, streamed from a laptop SSD

S-MoE (Seismic Mixture of Experts) is an open-source inference engine for fine-grained Mixture-of-Experts LLMs on Apple Silicon. It streams experts from the NVMe SSD at the moment they are needed, so the model never has to fit in RAM. Qwen3-235B runs on a MacBook from 32 GB of unified memory up.

I am a self-taught web developer. I built this with an AI agent that knows C++ and low-level programming, because I do not.


The story

A question I was not qualified to ask

To run a frontier model like Qwen3-235B, the usual answer is that all 235 billion parameters must sit in memory at once. At bfloat16 that is roughly 470 GB. A standard MacBook has 16 to 48 GB. The conclusion everyone draws is: rent a cloud GPU.

One detail kept nagging me. In a Mixture-of-Experts model, more than 95% of those parameters are silent at any given moment. At each layer, for each token, the router picks 8 experts out of 128. The rest is cold stone, and standard runtimes hold it in RAM "just in case".

So I asked the obvious, naive question: what if you only loaded what fires, just before it fires?

I did not know enough to know it was supposed to be impossible. That turned out to be useful.

The satellite man who looked sideways

The mental model came from a place far from AI. In 2022 Ing. Filippo Biondi published a paper in Remote Sensing on imaging the inside of the Great Pyramid of Giza with Synthetic Aperture Radar. Radar cannot penetrate rock. His idea was to stop fighting that: the pulse strikes the surface, the stone vibrates, and the vibrations carry the shape of what is inside. You measure what the barrier does to the signal.

I borrowed that intuition, and only the intuition:

Seismic concept S-MoE equivalent
Deep rock strata The expert weights, cold on the NVMe SSD
EM surface pulse The dense backbone running the current token
Generated phonons The router gates' exact activation echo
Acoustic receivers The async I/O streamer and its retained ring
Subsurface image The generated token

That is where the name comes from, and why the project speaks of mountains, strikes and vaults.

Built with an AI agent, honestly

I described the shape of what I wanted and asked naive questions. The agent wrote C++ and Metal. I read it, pushed back when it got too complex, and we iterated. I brought the curiosity and the direction; the agent brought the implementation depth. I understand a little more of the low-level code with each session.

The first target was a small model, DeepSeek-MoE-16B. Then the engine was refactored to be model-agnostic, and the real target arrived: Qwen3-235B, a 470 GB download onto a laptop.

What the numbers corrected

The project is measurement-driven, and two of my own ideas did not survive measurement.

  • The 16 GB claim. The original promise was "16 GB of RAM gives you the same intelligence as 512 GB, just slower". The experts really do occupy zero bytes of RAM, but the dense backbone of a 235B model weighs about 16 GB by itself. The honest floor for the frontier is 32 GB. A 16 GB Mac runs the same architecture with smaller MoE models. The manifesto was amended in public.
  • The Scout. The first design had a "Surface Scout" that predicted which experts the model would need next. Measured as a predictor, it covered 51.5% of the next experts at about 265 ms per token. Simply keeping recent experts in the ring covered 46.4% at zero cost, and the model's own router gate gives the exact answer with one small matrix multiply. The predictor was removed. The engine became faster and more faithful to the model at the same time.

The whole journey is written up as an eight-part build diary, linked at the bottom.


How it works

A model is prepared once, then served by three parts that run concurrently.

0. The shatter (offline, once)

shatter_moe.py takes a MoE checkpoint and splits it into two files:

  • The vault (.smoe): every routed expert, quantised to 4 bits and aligned to Apple Silicon's 16 KB page boundary so it can be read with Direct I/O. For Qwen3-235B this is about 117 GB.
  • The backbone (.scout.safetensors): embeddings, attention, norms and routing gates, kept in bfloat16. About 16 GB for the 235B.

The vault is self-describing. RoPE theta, routing top-k, attention geometry and lineage are written into an arch block from the checkpoint's own config.json, so the engine reads what it is given and needs no per-model configuration.

1. The resident backbone (in memory)

The backbone stays in unified memory. At every layer of every token it evaluates the real router gate on the true hidden state. That names the exact experts the model wants, the same ones it was trained to use.

2. The deep strata (on the SSD)

The vault rests cold on NVMe and holds zero bytes of RAM. Nothing is memory-mapped "just in case".

3. The acoustic receiver (in between)

Background workers read the named experts with Direct I/O (F_NOCACHE), bypassing the OS page cache, into a lock-free ring buffer a fraction of a second before the Metal GPU kernel runs them. Weights are dequantised directly in GPU registers.

The ring is also an LRU cache. Adjacent tokens reuse 46.4% of their experts, and those are already there, so they are never read again.

prompt token
    │
    ▼
backbone in RAM ──► router gate ──► "layer 37 needs experts 4, 19, 52, …"
                                         │
                       in the ring? ─────┤
                        yes: cache hit   │ no: Direct I/O read from the vault
                                         ▼
                                  ring buffer (unified memory)
                                         │
                                         ▼
                                Metal kernel ──► next layer ──► token

The rules the engine never breaks

  • Zero runtime heap allocations. malloc, new and std::vector::resize are illegal inside the token loop. Every buffer is carved at startup.
  • Direct I/O only. SSD to unified memory to GPU, with no page cache and no extra copies.
  • Atomic synchronisation only. No mutexes and no condition variables. The I/O thread and the GPU thread cannot block each other.
  • Decoupled routing. The gates measure, the streamer moves bytes, the kernel multiplies. None of them knows what the others do.

Tech stack

  • Core: C++20, data-oriented, POSIX primitives
  • Compute: Metal shaders behind an Objective-C++ bridge, zero-copy in unified memory
  • Tooling: Python for the shatter script, the chat console and the HTTP server

Key features

  • A frontier model on a laptop. Qwen3-235B (94 MoE layers × 128 experts) on a 32 to 48 GB MacBook, with about 16 GB resident and the ~117 GB vault at zero bytes of RAM.
  • Intelligence does not degrade, only speed does. A 32 GB Mac and a 512 GB Mac produce identical outputs from the same vault. More memory means a larger ring and more tokens per second.
  • Model-agnostic. Any compatible fine-grained MoE becomes a vault with one shatter command. No recompilation.
  • A model fleet, switchable at runtime. Every vault on disk is listed on /v1/models. Naming another one in the standard model field switches the engine.
  • OpenAI and Anthropic compatible. One process, one port, both protocols, streaming or not. Open WebUI, LibreChat, the openai and anthropic SDKs and Claude Code connect without being modified.
  • Built-in web chat. A dependency-free console at GET / with live turn timings, context and RAM meters, and a model dropdown.
  • Persistent sessions. The KV-cache and the ring survive across turns, so each message prefills only its new tokens. Follow-up turns answer in about 14 s however long the conversation is.
  • Layer-major batched prefill. Each layer's experts are read from the SSD once per prompt chunk, not once per token.
  • Private and offline. No cluster, no API key, nothing leaves the machine.
  • Open source, MIT. The code, the documentation and the mistakes are public.

The numbers

Measured on a standard 48 GB Apple Silicon MacBook, Q4 vault, 45-token prompt.

Qwen3-235B-A22B Qwen3-30B-A3B
Role The frontier The daily driver
Vault (Q4) ~117 GB, streamed from SSD ~14 GB, almost entirely in the ring
Decode 1.84 tok/s ~13.7 tok/s

For the 235B, since the latency work began:

Metric Before Now
Cold time to first token 92.7 s ~26 s
Decode speed 0.48 tok/s 1.84 tok/s
Warm follow-up turn grew with the conversation ~14 s, flat

I will not pretend 1.84 tokens per second is fast. It is a 235B model answering at reading speed on a machine you can close and put in a bag. The 30B is the one for daily use.


Try it

You need an Apple Silicon Mac and a fast internal SSD.

git clone https://github.com/melasistema/s-moe.gitcd s-moemake setup

Then download a supported checkpoint, shatter it and serve it:

make probe   MODEL=./checkpoints/qwen3-235b-instruct              # dry run, maps the topologymake shatter MODEL=./checkpoints/qwen3-235b-instruct OUT=./vault  # builds the vaultmake all                                                          # builds the engine.venv/bin/python serve.py --port 8000                             # OpenAI + Anthropic + web chat

Open http://127.0.0.1:8000/ and start typing, or point an existing client at it:

from openai import OpenAI  
client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="unused")  
ANTHROPIC_BASE_URL=http://127.0.0.1:8000 ANTHROPIC_API_KEY=unused claude  

The full guide, including the Qwen3-30B path for smaller machines, is in the documentation.


What is next

  • Q2 vaults. Decode is bound by SSD bandwidth. Halving every expert halves the reads and doubles the ring.
  • The 477B. GLM MoE 477B is the next summit. It is not shattered yet.
  • Paged KV-cache, so long conversations do not depend on RAM.
  • Adaptive quantisation, keeping popular experts at 4 bits and compressing niche ones to 2.

S-MoE is an experiment. It is incomplete, and I would rather show the measurements than oversell it.


The build diary

  1. Rock, Paper, Silicon
  2. "Hello, World! "
  3. The Mountain Learns to Converse
  4. The End of Prophecy
  5. The Mountain Predicts Itself
  6. The Mountain Runs
  7. S-MoE Started to Speak OpenAI
  8. The Mountain Exhales

The seismic idea is a philosophical translation of the published work of Ing. Filippo Biondi on SAR-based subsurface imaging. The method is his; I only borrowed the intuition.

Speed is a privilege. Intelligence is not.

🔥 Join developers growing publicly
Share your knowledge, build in public, and grow your developer presence with a global community.

More Posts

Sovereign Intelligence: The Complete 25,000 Word Blueprint (Download)

Pocket Portfolio - Apr 1

The Sovereign Vault — A Comprehensive Guide to Protocol-Driven AI

Ken W. Algerverified - Jun 4

Architecting a Local-First Hybrid RAG for Finance

Pocket Portfolio - Feb 25

S-MoE Started to Speak OpenAI — Point Any Chat App at Your MacBook and Watch a 235B Model Think

melasistema - Jul 20

The Mountain Exhales — Riding the 30B Wind Before the 477B Avalanche

melasistema - Jul 21
chevron_left
3.9k Points • 88 Badges
Bolzano - Italy • github.com/melasistema
28Posts
12Comments
16Connections
Full Stack Developer and Technology Consultant with a solid ten-year experience in supporting startu... Show more

Related Jobs

View all jobs →

Commenters (This Week)

4 comments
1 comment

Contribute meaningful comments to climb the leaderboard and earn badges!