Today, we'd love to announce Nativ's Day-0 support for GLM-5.3-Flash, powered by mlx-vlm. GLM-5.3-Flash is built on the GLM-5-Next architecture and runs on Nativ out of the box.

As a 320-billion-parameter sparse MoE — and a built-in vision-language model — it pairs 288 experts (top-8, plus one shared expert) to activate ~16B parameters per token. Its 45-layer hybrid backbone applies DeepSeek Sparse Attention at every 4th layer while the remaining 34 layers run Kimi-delta linear attention, carrying mHC hyper-connection residual streams across a 1M-token context. As a frontier open-weight model, running it in MLX takes a high-memory Mac: a 256 GB Mac Studio (M3 Ultra) for the 4-bit / MXFP4 builds, or a single 128 GB Mac for the 2-bit-expert version.

Earlier this year we got GLM-5, GLM-5.2, and GLM-5.3 — and now GLM-5.3-Flash. Safe to say, the people at Z.ai have long been busy pushing the frontier of open source.

This 320B-sized behemoth — while not the boldest or biggest we've seen from Z.ai — certainly carries weight, and packs a punch. With the end of dense models, it's one of many recent sparse MoEs the open-source community has been blessed with. Unsurprisingly, like many of its predecessors, GLM-5-Next builds on open research spanning labs and research groups across academia and industry.

TL;DR

  • Day-0 support: GLM-5.3-Flash reuses the GLM-5-Next architecture and runs on mlx-vlm from day one, with specific optimizations across prefill and decode.
  • MTP support: day-0 support with the MTP draft model released at zai-org/GLM-5.3-Flash and zai-org/GLM-5.3-Flash-BF16.

/ Architecture

GLM-5-Next's architecture, and how Nativ serves it.

New approach. GLM-5-Next takes a bolder, different tack — arguably a design much closer to the recently released Kimi K3. For its size — about half of the previous GLM-5 and 5.2 — it still holds 288 experts, roughly 32 more than either predecessor. Like a mini Kimi K3, GLM moves from pure sparse to a hybrid-linear recipe, becoming the first to pair linear with sparse attention.

/ Performance optimizations

Performance optimizations.

Parallel linear attention

The 34 Kimi-delta layers are the bulk of the model, and their inherently sequential nature (token i depends on token i−1) is obviously not ideal for our parallel, throughput-hungry GPUs. The chunked-prefill scan naturally follows a design similar to FlashKDA, expressed through our gated_delta_update:

out, state = gated_delta_update(
            q,
            k,
            v,
            a,
            b_o,
            fg.A_log.reshape(self.num_heads, 1),
            fg.dt_bias.reshape(self.num_heads, self.head_dim),
            state=state,
            lower_bound=fg.safe_gate_lower_bound,
)

Memory-safe sparse indexer

Sparse attention, pioneered by DeepSeek, avoids a full lookup between every token and all of its predecessors by employing a lightning indexer that scores an earlier position s for the current position t as:

$$ I_{t,s} = \sum_{j=1}^{H^{I}} w_{t,j}\, \mathrm{ReLU}\!\left(q_{t,j} \cdot k_{s}\right). $$

With its own k and v projections, it compresses the history into small pools (groups of a few tokens squashed together), then scores the current query against those pools to rank them.

In our preliminary testing, a 64K context was enough to force mlx-vlm to allocate 137 GB and essentially crash. Since the indexer must score all S queries against all S/4 key-pools across its 32 scoring heads to rank them, the score matrix grows quadratically with context — a 64K prefill materializes 65,536 × 32 × 16,384 × 4 bytes ≈ 137 GB at once. Sound familiar? That's because it mirrors softmax attention's $O(S^2)$ memory. Mitigating it requires knowing that each top-k is independent, so we never need all S queries' scores resident together. Dropping the query axis yields $O(\text{chunk} \cdot P)$ — linear in S — about 1 GB per chunk, for a max of 6 GB, drastically lower than 137 GB.

/ Batching

Maximizing & distributing intelligence.

With the recent release of the M6 and M5 Pro, saturating the GPU has never been more important — and under a single-batch scenario, that doesn't happen. Static single-batching does not use the GPU efficiently: for N requests it waits until the entire batch is finished, wasting compute that could have run in the meantime. With BatchGenerator in mlx-vlm, we run dynamically managed batches — new requests admitted and finished ones evicted mid-flight — so each expensive weight read is amortized across every active sequence. That easily took us from 34 tok/s single-stream to 191 tok/s at 32 concurrent requests — almost 5.5× — with prefill up 3.3×, and peak memory barely moving (179 → 188 GB).