On an Apple M5 Max, our implementation reaches 817 text embeddings per second with batches of 32 short texts and closely matches the supplied reference model across all 20 documentation examples. Below, we walk through the architecture, performance across modalities, and the trade-offs of quantization. The implementation is available in MLX-VLM.

One space for every modality

EmbeddingGemma 2 produces a 768-dimensional vector for each input. That input can contain a single modality or a combination: a photograph with a description, several images, text interleaved with audio, or video frames with accompanying text. The processor preserves content order, allowing the model to represent the combined input as one embedding.

This shared representation supports cross-modal retrieval, semantic similarity, clustering, and classification pipelines. For example, a text query can be compared with embeddings of images, recordings, and video clips in the same search index. Optional task prompts provide context for retrieval, classification, clustering, code search, and other uses. Their effectiveness should be evaluated on the application’s own data. Model features and usage

Choose your embedding size

Matryoshka Representation Learning adds flexibility to the output size. Applications can retain the first 512, 256, or 128 dimensions of an embedding and normalize the result again. Queries and indexed documents must use the same dimension.

At the same storage precision, a 256-dimensional vector uses one-third of the raw vector storage of a 768-dimensional vector; 128 dimensions use one-sixth. This reduces storage and distance-comparison costs. The encoder still performs its forward pass, and retrieval quality at each dimension remains something to measure on the target dataset.

Inside the architecture

The architecture combines Gemma 4 vision and audio encoders with a 24-layer bidirectional text encoder. Images and sampled video frames pass through the vision encoder, while audio passes through the audio encoder. Their features are projected into the text encoder’s 512-dimensional input space and inserted alongside text-token embeddings.

EmbeddingGemma 2 architecture, with data flowing from bottom to top: text embeddings and projected vision and audio features enter a 24-layer bidirectional encoder, followed by projection, mean pooling, and L2 normalization to a 768-dimensional embedding. Detail panels show gated GELU feed-forward, local and full attention, and per-layer input modules.
EmbeddingGemma 2 architecture. Data flows from bottom to top through the shared encoder to a normalized embedding. View full-size diagram

The shared encoder alternates five local-attention layers with a full-attention layer. Local attention uses a radius of 512 positions in each direction; full attention connects information across the complete input. Because attention is bidirectional, token representations can incorporate context from both earlier and later positions.

Each layer also receives a learned projection of the original input representations through the model’s per-layer input mechanism. After the encoder, a linear projection maps token states to 768 dimensions. Mean pooling combines the non-padding token representations, and float32 L2 normalization produces the final vector. Prompt and media tokens participate in pooling.

The complete model contains approximately 744 million parameters. Disabling the audio tower reduces this to approximately 439 million; disabling both media towers leaves a text-only model of approximately 271 million parameters. This allows deployments to load only the modalities they need.

Visual processing is configurable as well. Supported soft-token budgets range from 70 to 1,120 per image or frame. Video processing exposes sampling rate, frame limits, overflow handling, and optional timestamps. These controls affect how much visual information enters the encoder and therefore the computation required.

Performance on the M5 Max

Our performance measurements used an Apple M5 Max with 128 GiB of unified memory. The batch sweep covered sizes 1, 2, 4, 8, 16, and 32 across BF16 and three quantized formats.

The following table shows BF16 throughput and peak active memory:

BF16 throughput and peak active memory
InputBatch 1Batch 32Throughput changeMemory at batch 32
Text, 128 tokens167.3 texts/s817.4 texts/s4.9×2.51 GB
Image22.8 images/s21.2 images/s0.93×14.44 GB
Audio, 5.855-second recording38.5 recordings/s85.4 recordings/s2.2×3.53 GB
Video, 5-second clip sampled into five frames9.75 videos/s8.80 videos/s0.90×33.66 GB

A batch of 32 short texts completed in approximately 0.039 seconds. Longer text reached its best measured BF16 throughput at a smaller batch size: 512-token inputs peaked at approximately 190 texts per second with batch 8.

Audio also benefited from batching, reaching its best measured BF16 throughput at batch 16. Image and video workloads behaved differently. Their throughput stayed roughly flat or declined, while memory consumption increased. For video, peak active memory grew from approximately 3.10 GB at batch 1 to 33.66 GB at batch 32. Detailed batch results

Batched throughput for text, image, audio, and video across BF16, 8-bit, 6-bit, and 4-bit precision on an Apple M5 Max. Text and audio throughput rise with batch size, while image and video throughput remain roughly flat.
EmbeddingGemma 2 batched throughput on MLX-VLM. View full-size chart

Quantization and fidelity

Quantization primarily reduced weight storage in these tests. The standard conversion policy quantized the text encoder and audio projection while retaining the vision encoder, audio encoder, and vision projection in BF16. Consequently, whole-model savings were smaller than the savings within the text encoder.

Weight storage and embedding fidelity by precision
FormatTotal weight storageStorage reductionMinimum embedding cosine vs. FP32 reference
BF161.489 GB—0.999791
8-bit1.234 GB17.1%0.999662
6-bit1.166 GB21.7%0.997708
4-bit1.098 GB26.2%0.967160

8-bit provided a useful memory–fidelity trade-off, closely preserving the reference embeddings while reducing total weight storage by approximately 17%. Both 8-bit and 6-bit passed all 20 examples under the quantization validation thresholds.

Standard affine 4-bit quantization introduced substantially more numerical drift, failing those thresholds in 18 of the 20 examples. Its outputs remained finite and normalized, but the additional drift makes downstream retrieval evaluation particularly important.

Quantization did not produce a consistent speed improvement. BF16 was fastest in the single-input text, image, and audio measurements; small video differences were within observed timing variation. Quantization and latency measurements

Validation and methodology

Numerical validation covered all 20 supplied documentation examples, with 32 paired forward comparisons per precision and additional checks of the shorter Matryoshka vectors. Float32 MLX embeddings had a maximum absolute component error of approximately 5.79 × 10⁻⁷ against the float32 reference. BF16 achieved a minimum embedding cosine similarity of 0.9997906.

The batch sweep also passed all 240 batch-versus-single consistency checks, with a minimum cosine similarity of approximately 0.999886 against the same checkpoint’s individual output. These measurements establish implementation fidelity and batching consistency; they do not establish retrieval quality on a benchmark such as MTEB.

All performance figures measure warmed, synchronized model execution, including the applicable media encoders, pooling, and normalization. They exclude file decoding, preprocessing, model loading, and initial compilation. The batch tests repeat fixed inputs and therefore describe homogeneous workloads rather than a production queue containing varying lengths and media sizes. Memory figures represent active MLX allocations, not total process memory.

Putting it into practice

For the tested M5 Max, BF16 provides a strong starting point when speed and numerical fidelity matter most. Larger batches are useful for short-text indexing and audio workloads, while image and video applications should choose batch sizes carefully. Applications constrained by memory can first disable unused modality towers, then evaluate 8-bit quantization and smaller embedding dimensions against their own retrieval requirements.

Back to the blogNATIV / ENGINEERING