Today, we'd love to announce Nativ's Day-0 support for GLM-5.3-Flash, powered by mlx-vlm. GLM-5.3-Flash is built on the GLM-5-Next architecture and runs on Nativ out of the box.
As a 320-billion-parameter sparse MoE — and a built-in vision-language model — it pairs 288 experts (top-8, plus one shared expert) to activate ~16B parameters per token. Its 45-layer hybrid backbone applies DeepSeek Sparse Attention at every 4th layer while the remaining 34 layers run Kimi-delta linear attention, carrying mHC hyper-connection residual streams across a 1M-token context. As a frontier open-weight model, running it in MLX takes a high-memory Mac: a 256 GB Mac Studio (M3 Ultra) for the 4-bit / MXFP4 builds, or a single 128 GB Mac for the 2-bit-expert version.
Earlier this year we got GLM-5, GLM-5.2, and GLM-5.3 — and now GLM-5.3-Flash. Safe to say, the people at Z.ai have long been busy pushing the frontier of open source.
This 320B-sized behemoth — while not the boldest or biggest we've seen from Z.ai — certainly carries weight, and packs a punch. With the end of dense models, it's one of many recent sparse MoEs the open-source community has been blessed with. Unsurprisingly, like many of its predecessors, GLM-5-Next builds on open research spanning labs and research groups across academia and industry.
TL;DR
- Day-0 support: GLM-5.3-Flash reuses the GLM-5-Next architecture and runs on mlx-vlm from day one, with specific optimizations across prefill and decode.
- MTP support: day-0 support with the MTP draft model released at zai-org/GLM-5.3-Flash and zai-org/GLM-5.3-Flash-BF16.
/ Architecture
GLM-5-Next's architecture, and how Nativ serves it.
New approach. GLM-5-Next takes a bolder, different tack — arguably a design much closer to the recently released Kimi K3. For its size — about half of the previous GLM-5 and 5.2 — it still holds 288 experts, roughly 32 more than either predecessor. Like a mini Kimi K3, GLM moves from pure sparse to a hybrid-linear recipe, becoming the first to pair linear with sparse attention.
/ Performance optimizations
Performance optimizations.
Parallel linear attention
The 34 Kimi-delta layers are the bulk of the model, and their
inherently sequential nature (token i depends on token
i−1) is obviously not ideal for our parallel, throughput-hungry
GPUs. The chunked-prefill scan naturally follows a design similar to
FlashKDA, expressed through our gated_delta_update:
out, state = gated_delta_update(
q,
k,
v,
a,
b_o,
fg.A_log.reshape(self.num_heads, 1),
fg.dt_bias.reshape(self.num_heads, self.head_dim),
state=state,
lower_bound=fg.safe_gate_lower_bound,
)
Memory-safe sparse indexer
Sparse attention, pioneered by DeepSeek, avoids a full lookup between every token and all of its predecessors by employing a lightning indexer that scores an earlier position s for the current position t as:
$$ I_{t,s} = \sum_{j=1}^{H^{I}} w_{t,j}\, \mathrm{ReLU}\!\left(q_{t,j} \cdot k_{s}\right). $$
With its own k and v projections, it compresses the history into small pools (groups of a few tokens squashed together), then scores the current query against those pools to rank them.
In our preliminary testing, a 64K context was enough to force mlx-vlm to allocate 137 GB and essentially crash. Since the indexer must score all S queries against all S/4 key-pools across its 32 scoring heads to rank them, the score matrix grows quadratically with context — a 64K prefill materializes 65,536 × 32 × 16,384 × 4 bytes ≈ 137 GB at once. Sound familiar? That's because it mirrors softmax attention's $O(S^2)$ memory. Mitigating it requires knowing that each top-k is independent, so we never need all S queries' scores resident together. Dropping the query axis yields $O(\text{chunk} \cdot P)$ — linear in S — about 1 GB per chunk, for a max of 6 GB, drastically lower than 137 GB.
/ Batching
Maximizing & distributing intelligence.
With the recent release of the M6 and M5 Pro, saturating the GPU has
never been more important — and under a single-batch scenario, that
doesn't happen. Static single-batching does not use the GPU
efficiently: for N requests it waits until the entire batch
is finished, wasting compute that could have run in the meantime. With
BatchGenerator in mlx-vlm, we run dynamically managed
batches — new requests admitted and finished ones evicted mid-flight —
so each expensive weight read is amortized across every active
sequence. That easily took us from 34 tok/s single-stream to
191 tok/s at 32 concurrent requests — almost 5.5× —
with prefill up 3.3×, and peak memory barely moving (179 → 188 GB).