On-device inference often runs at batch size one. One application decodes one active sequence on one device.
This is quite different from a busy hosted service where requests from many users reach the same accelerator pool. The runtime can schedule several sequences in one batch. A single model pass then advances multiple requests. This improves total throughput.
An on-device application usually has no such cross-user request pool. Each installed copy handles its own prompts. A coding agent may invoke the model many times. However, later calls often depend on earlier outputs. Those calls usually arrive in sequence and cannot form a useful batch.
For this workload, one-sequence latency becomes the main target. During ordinary decoding, each token requires another model pass.
At batch size one, the runtime cannot share that weight traffic across user sequences. Quantization reduces the bytes moved during each pass. Speculative decoding tries to accept several tokens from one pass.
Losing cross-user batching is a constraint. Yet it also narrows the system’s problem. The local engine does not route requests across a fleet. It does not schedule work from many tenants. It can optimize for one active decode on known hardware.
Uzu is an open-source implementation of that approach. When an application uses Uzu, the engine and model run on the user’s device, fully locally.
A generic local runtime can load a quantized checkpoint. Uzu builds both the checkpoint and runtime. It designs the quantization, kernels, and speculative path as one system.
On a 14-inch MacBook Pro with a base M5 chip, the same 4-bit Qwen3.5 9B model wrote a Python function at 92 to 117 tokens per second with Uzu, across repeated runs of the same prompt. In our side-by-side recording, llama.cpp and MLX managed only 22 and 25.
In this article, we will trace where that gap comes from, through quantization and speculative decoding. Then we will run Uzu on that M5, inspect its metrics, and compare it with llama.cpp and MLX on the same prompt.
Why batch-one decoding is memory-bound
LLM inference has two distinct phases.
Prefill processes the prompt tokens in parallel.
Decode generates the response one token at a time.
Decode is the relevant phase where each new token requires another model pass. During that pass, the runtime reads weight blocks and computes the next-token scores.
At batch size one, a forward pass advances only one sequence.
The overall arithmetic is small relative to the weight data moved. To produce one token, the chip streams every weight of the model from memory into the GPU once and does roughly one multiply-add with each.
The GPU can finish those calculations before memory brings the next weight block. So the computation pace is often bottlenecked by the memory bandwidth.
This gives a hard limit. A base M5 moves 153 GB of data per second from memory. A 4-bit Qwen3.5 9B checkpoint weighs about 5.2 GB. If every pass produces one token, the model can write at most 153 ÷ 5.2 ≈ 29 tokens per second on that chip, however fast the GPU computes.
Real engines land a little below that ceiling, because each pass also reads other data, such as the cache of earlier tokens.
Quantization is a standard technique to improve this, whereby using smaller weights reduces the bytes read during each pass.
However, quantization still optimizes the same one-token loop. The main model must run again for every generated token.
In fact, quantization is already quite common, so now let’s understand what makes Uzu different as an inference engine.
How Uzu optimizes bytes and tokens per pass
Uzu targets two quantities in the decoding loop. It reduces bytes transferred per pass and increases accepted tokens per pass.
Quantization handles the first quantity. It stores model weights with fewer bits. Smaller checkpoints require less memory traffic and leave more unified memory for context and application state.
Speculative decoding handles the second quantity. It verifies several candidate tokens together. If several candidates are approved, the sequence advances by several tokens.
Many local engines already support 4-bit weights. Uzu’s distinction lies in the hardware path around that format.
Mirai publishes multiple optimized checkpoints for the same base model. It labels them by size tier. M means Medium, while L means Large.
These labels describe the checkpoint, not a different model architecture.
The Medium checkpoint uses 4-bit asymmetric integer quantization.
The Large checkpoint uses 8-bit symmetric quantization.
Medium favors a smaller memory footprint, while Large keeps more bits per weight. Mirai’s quantization write-up describes how it combines post-training quantization with quantization-aware distillation.
Using the integer format is intentional since it helps both the hardware and the compression.
More specifically, unlike older Apple GPUs, the M5 reaches its highest arithmetic throughput in int8.
Uzu therefore also quantizes activations, the intermediate values passed between layers, to 8 bits on the fly. With both sides of each matrix multiplication stored as integers, the speculative verification step runs on the M5’s hardware-accelerated int8 matrix path.
Lower precision creates another problem.
More specifically, group-wise quantization introduces rounding error when a group contains an outlier.
Every value in the group shares one scale. That scale must represent the largest magnitude. Smaller values then use fewer integer levels and incur more rounding error.
Uzu applies a block-diagonal Random Hadamard Transform before quantization. A Hadamard matrix contains only +1 and -1 entries, plus one overall scale.
Multiplying a block of weights by it mixes every value into every position, so a single large outlier turns into many moderate values.
“Random” refers to random sign flips applied before the mixing, and “block-diagonal” means each block of 32 values is mixed on its own. The transform is exactly reversible, so apart from rounding it does not change what the model computes.
Mirai selected 32 elements to match the Apple GPU kernel’s SIMD width. SIMD means that one instruction runs across 32 GPU threads at once, so one group of threads can transform one block.
The transform block and quantization group control different operations.
The transform block defines which values the Hadamard matrix mixes.
The quantization group defines which values share quantization parameters.
These sizes do not need to match.
Uzu’s differentiation is the joint design of the number format, transform, and runtime kernels. Quantization reduces memory traffic and computation within each model pass.
However, ordinary decoding still returns one token per target-model pass. Speculative decoding is the next lever it utilizes to improve that bottleneck.
Producing several tokens from one model pass
Speculative decoding adds a small draft model beside the main model.
The draft (small) model proposes future tokens cheaply. The main model checks those positions together. It accepts the valid prefix and resumes from the first rejected token.
Suppose the drafter proposes five tokens and the main model agrees with the first four. That one verification pass advances the sequence by five tokens, the four accepted drafts plus the main model’s own token at the fifth position.
Ordinary decoding would need five separate main-model passes for the same text.
The main model still controls the final output. The drafter only proposes tokens that can be checked in parallel. A correct acceptance procedure preserves the target model’s output distribution.
The verification step requires more computation than one-token decoding.
But the path is still preferred because local batch-one inference often has arithmetic capacity sitting idle while weights move through memory.
Speculative decoding spends that spare arithmetic to reduce the number of main-model passes.
The trade-off depends on acceptance. Longer drafts create larger verification batches. They also create more chances for an early mistake to invalidate later tokens.
Many model-native systems draft three or four tokens. A common example is multi-token prediction (MTP), where extra prediction heads trained into the model itself guess the next few tokens.
Uzu uses budgets of 16 to 32 tokens on its M5 path. The Qwen3.5 9B checkpoint used later in this article ships with a 16-token tree budget for M5-series chips.
Those larger batches give the M5’s Neural Accelerators enough work. These are matrix units built into each GPU core, and they only pay off when a matrix multiplication has many rows. Checking 16 candidate tokens at once gives them 16 rows of work instead of 1.
Longer parallel drafts introduce a coherence problem.
Issue with parallel draft generation
A fast parallel drafter predicts several future positions at once.
As expected, each position sees the existing context, but not the tokens selected at earlier draft positions.
So consider a prompt ending with “She went to the store to buy.” Independent positions may prefer “a,” “few,” “of,” and “milk.” Each token can look plausible alone. Together they produce “a few of milk.”
An early mismatch rejects the remaining chain. This is why acceptance tends to weaken as a parallel draft grows. More positions create more opportunities for individually plausible tokens to form an invalid sequence.
The Trees from Marginals paper introduces Weaver to repair this weakness. Weaver is a 56.7M-parameter autoregressive adapter in the paper’s research setup. It operates on candidate shortlists rather than projecting across the full vocabulary.
A parallel DFlash drafter first predicts likely candidates for each future position in one go. Weaver then selects from those candidates in sequence. Every selection can depend on earlier selections.
The earlier example can now become “a gallon of milk.” The later choices follow the path already selected.
Weaver can also retain several coherent paths. One branch might contain “a gallon of milk.” Another might contain “a bottle of juice.” The main model verifies the tree and accepts the longest valid path.
The tree shape can adapt to the prompt. Predictable code can support a long, narrow branch. Open-ended writing may need shorter branches with more alternatives.
The paper reports a 24.7 percent improvement over an optimized DFlash baseline. It also reports a 4.37-fold speedup over autoregressive decoding. Those experiments used CUDA and SGLang. They are research results, not universal Uzu measurements.
Uzu applies the same drafting idea to its Apple-focused runtime. The remaining problem is verifying trees efficiently on the target model.
Verifying a draft tree without restoring model state
Some model layers keep a running internal state. Each new token reads that state and updates it. With one sequence, the runtime applies these updates in order.
A draft tree creates several possible next sequences. Each branch would update the state differently. If the runtime updates its live state while checking one branch, it must restore the earlier state before checking another. Repeating that process across a large tree adds memory operations and duplicate work.
Qwen3.5 and Qwen3.6 contain stateful layers called Gated DeltaNet layers. An attention layer keeps a growing cache of every past token. A Gated DeltaNet layer instead compresses everything it has seen into a fixed-size state, and every accepted token changes that state.
In Qwen3.5 9B, 24 of the 32 layers are of this type.
Uzu separates checking a branch from committing its state. During verification, it computes the values needed by every candidate branch without changing the live state. The main model then selects the accepted path. Uzu applies state updates only for the accepted tokens.
The Uzu implementation note calls this rollback-free tree verification. Its Metal kernels compute the DeltaNet outputs for the whole tree in a single recurrence-free pass, instead of stepping through each branch token by token.
State is then committed only along the chosen path, so there is never a wrong update to roll back.
This matters because speculative decoding only helps when verification costs less than ordinary token-by-token decoding. Repeated state restoration would consume part of that gain.
Together, Weaver and rollback-free verification let Uzu propose long token sequences and check them without repeatedly restoring model state.
What this looks like on a base M5
To see the difference between ordinary decoding, conventional speculative decoding and Uzu’s approach, we ran the same Qwen3.5 9B model three ways on a MacBook Pro with a base M5 chip and 16 GB of unified memory:
llama.cpp with no speculative decoding
llama.cpp with the model’s built-in MTP heads drafting 3 tokens ahead, the conventional setup
Uzu with DFlash and Weaver
All three used 4-bit weights and the same prompt: “Write a python function to merge two sorted lists.”
Plain llama.cpp lands just under the 29 tokens-per-second limit from earlier. Three-token MTP drafts push it past that limit, to about 1.6 times plain decoding. In the recorded run, Uzu’s verifier accepted 7.5 tokens per target-model pass on average, which is how it reaches more than three times the one-token limit on the same chip.
Published benchmark comparisons
There are already some published benchmarks for Uzu that measure output speed, input speed, memory, and quantization quality.
The following speculative result uses Qwen3.6 27B Mirai-M. The checkpoint occupies 15.6 GB. Mirai measured it on an M5 Max with 128 GB of memory.
These figures put Uzu at
2.1 times MTPLX
3.8 times llama.cpp
4.4 times MLX.
Moreover, Uzu reaches 35.54 tokens per second without speculation. Its speculative result reaches 114 tokens per second. This provides a 3.2-times gain from the speculative path.
There are also results on output generation speed by prompt type.
Code and maths contain many predictable continuations. Open-ended conversation gives the drafter fewer reliable tokens. Speculative speed is therefore a workload property, not a fixed model number.
Run a model locally and inspect the runtime
The fastest way to check Uzu is its command-line tool. It installs through Homebrew and downloads the selected model on first use.
brew install mirai
mirai --model trymirai/Qwen3.5-9B-M \
-m "Write a python function to merge two sorted lists"We ran this with Mirai CLI 0.6.1 on a MacBook Pro with a base M5 chip and 16 GB of unified memory, on macOS 26.6.2. The first run downloads the model.
Here is a run in real time:
After the answer, the CLI printed these stats:
time to first token: 0.12 s
prefill speed: 161.11 t/s
generation speed: 98.10 t/s
tokens per forward pass: 8.39 t/f
memory used: 5.77 GB
total energy: 114.90 J
input energy per token: 0.11 J/tok
output energy per token: 0.16 J/tok
duration: 7.36 sHere is what these numbers mean:
Time to first token is how long you wait before the first token of the answer appears. Most of it goes into reading the prompt.
Prefill speed is how fast the engine reads the prompt, in prompt tokens per second. Our prompt is only about 20 tokens, which is too short for this number to say much. It climbs with longer prompts.
Generation speed is how fast the answer is written, in output tokens per second. This is the decode phase from earlier, the one held back by the memory limit.
Tokens per forward pass is how many output tokens each pass of the 9B model produced on average. Ordinary decoding gives exactly 1.00.
Memory used is the unified memory the engine took, mostly the model’s weights plus the drafter and the cache.
The 8.39 t/f value shows the speculative path at work. On average, each forward pass of the 9B model produced about 8 tokens. That is how generation reaches 98 tok/s on a chip whose one-token limit is about 29.
Speeds vary a little between runs. Across our runs of this command, generation speed ranged from 92 to 117 tok/s, with 7.5 to 9.6 tokens per pass.
Speculation is enabled per chip. This checkpoint’s speculator configuration lists the M5, M5 Pro and M5 Max, so on M1 to M4 Macs the same command runs without a drafter and reports 1.00 t/f.
The CLI exposes latency, throughput, memory, energy, and speculation data. We can now call the same runtime from an application.
Call Uzu through a local application endpoint
Uzu can expose an OpenAI-compatible HTTP endpoint. The model still runs on the user’s Mac. The HTTP request only crosses the local loopback interface.
Start the server in one terminal:
mirai server --model trymirai/Qwen3.5-9B-M --port 8000 --no-prefix-cacheThe server loads one local checkpoint and listens on 127.0.0.1, process on the same Mac can invoke it. This is useful for an existing application that already interacts via the OpenAI chat protocol.
By default, Uzu keeps a prefix cache. It stores the already-processed start of earlier prompts, so a repeated prompt can skip most of its prefill. --no-prefix-cache turns this off, so every test request does the full work and repeated runs stay comparable.
An existing application can switch to Uzu by pointing its OpenAI client at http://127.0.0.1:8000/v1. The next section uses one small client to time Uzu, MLX and llama.cpp in exactly the same way.
Use one client for Uzu, MLX, and llama.cpp
Here is the same prompt run through each engine on the base M5. Each engine ran on its own, with nothing else on the machine.
llama.cpp and MLX both sit under the 29 tokens-per-second limit, as expected when every pass produces one token. In this recording, Uzu writes about 3.7 times faster than MLX on the same chip.
To run this comparison with your own prompts, start each engine as a local server and time all three with one client.
The local server interface gives us a practical comparison method.
Uzu, MLX LM, and llama.cpp can all expose OpenAI-style chat endpoints. We can run Qwen3.5 9B through all three and send the same prompts to each server.
Start Uzu with Mirai’s 4-bit Medium checkpoint:
mirai server --model trymirai/Qwen3.5-9B-M --port 8000Start MLX LM with the 4-bit MLX conversion of the same base model:
brew install uv
uv tool install mlx-lm
uv tool update-shelluv tool update-shell adds the directory containing mlx_lm.server to your shell path.
Open a new terminal after running it, then start the server:
mlx_lm.server --model mlx-community/Qwen3.5-9B-4bit --port 8001Start llama.cpp with the Q4_K_S GGUF checkpoint used in Mirai’s published comparison:
brew install llama.cpp
llama-server --version
llama-server \
-hf unsloth/Qwen3.5-9B-MTP-GGUF:Q4_K_S \
--alias qwen3.5-9b-q4-k-s \
-ngl 99 -fa on -c 8192 -np 1 \
--port 8002-c 8192 keeps llama.cpp from reserving the model’s full 262K context, which would not fit next to the weights on a 16 GB Mac. -np 1 runs a single slot. To turn on llama.cpp’s MTP drafting, as in the first video, add --spec-type draft-mtp --spec-draft-n-max 3. MTP needs -np 1.
Each server downloads its checkpoint the first time it starts. Run the servers one at a time so that only one copy of Qwen3.5 9B occupies unified memory during a measurement.
The three checkpoints come from the same Qwen3.5 9B model, but they do not use the same file format or quantization recipe.
Uzu uses Mirai-M, MLX uses its own 4-bit conversion, and llama.cpp uses Q4_K_S. This is therefore a comparison of complete engine paths that a developer can run. It is not a kernel-only test.
Run the benchmark client against ports 8000, 8001, and 8002, one server at a time.
The harness should keep the prompt text, output limit, sampling settings, and run count fixed. Use comparable quantization quality, not only similar bit width. Record native engine metrics beside the end-to-end client time.
Finally, here’s a standalone client. Save it as compare_client.py. It uses only Python’s standard library and sends the same deterministic prompt to any OpenAI-compatible local endpoint.
It turns thinking off so every engine answers directly, and it asks each server for token usage. It reports time to first token, end-to-end time and decode speed.
Uzu also reports how many target-model passes it ran, so for Uzu the script prints tokens per pass too.
import json, os, sys, time, urllib.request
PROMPT = "Write a Python function that merges two sorted lists."
body = json.dumps({
"model": os.environ["MODEL"],
"messages": [{"role": "user", "content": PROMPT}],
"temperature": 0, # same answer every run
"max_tokens": 1024,
"stream": True, # tokens arrive as they are made
"stream_options": {"include_usage": True}, # server reports token counts
"chat_template_kwargs": {"enable_thinking": False}, # answer directly, no thinking
}).encode()
request = urllib.request.Request(os.environ["ENDPOINT"], body, {"Content-Type": "application/json"})
start = time.perf_counter()
first = last = None
usage = {}
with urllib.request.urlopen(request, timeout=600) as response:
for line in response:
line = line.decode().strip()
if not line.startswith("data:"):
continue
data = line[5:].strip()
if data == "[DONE]":
break
event = json.loads(data)
usage = event.get("usage") or usage
for choice in event.get("choices") or []:
delta = choice.get("delta") or {}
text = delta.get("content") or ""
if text or delta.get("reasoning_content") or delta.get("reasoning"):
last = time.perf_counter()
first = first or last
print(text, end="", flush=True)
if first is None:
sys.exit("No tokens received. Check the server log.")
tokens = usage.get("completion_tokens", 0)
print(f"\n\ntime to first token: {first - start:.2f} s")
print(f"end-to-end time: {time.perf_counter() - start:.2f} s")
if tokens > 1 and last > first:
print(f"decode speed: {(tokens - 1) / (last - first):.1f} tok/s")
if usage.get("spec_verify_ct"): # only Uzu reports how many model passes it ran
print(f"tokens per pass: {tokens / usage['spec_verify_ct']:.2f}")Run each engine, one at a time
Run only one server at a time, so a single copy of the model sits in unified memory and no two engines compete for memory bandwidth.
For each engine, start the server in one terminal, then run the client twice from a second terminal. The first run includes warm-up work, so discard it and keep the second. Stop the server with Ctrl-C before moving on.
For Uzu, start the server in the first terminal:
mirai server \
--model trymirai/Qwen3.5-9B-M \
--port 8000 \
--no-prefix-cacheRun the client from a second terminal:
ENDPOINT=http://127.0.0.1:8000/v1/chat/completions \
MODEL=trymirai/Qwen3.5-9B-M \
python3 compare_client.pyRun the client twice. Discard the first result because it includes warm-up work. Keep the second result and the decode statistics printed by the Uzu server. Then stop Uzu with Ctrl-C.
Start MLX in the first terminal:
mlx_lm.server \
--model mlx-community/Qwen3.5-9B-4bit \
--port 8001Run the same client from the second terminal:
ENDPOINT=http://127.0.0.1:8001/v1/chat/completions \
MODEL=mlx-community/Qwen3.5-9B-4bit \
python3 compare_client.pyAgain, discard the first run and keep the second. Stop MLX with Ctrl-C.
Start llama.cpp in the first terminal:
llama-server \
-hf unsloth/Qwen3.5-9B-MTP-GGUF:Q4_K_S \
--alias qwen3.5-9b-q4-k-s \
-ngl 99 -fa on -c 8192 -np 1 \
--port 8002Run the client from the second terminal:
ENDPOINT=http://127.0.0.1:8002/v1/chat/completions \
MODEL=qwen3.5-9b-q4-k-s \
python3 compare_client.pyDiscard the first run and keep the second. The comparison should keep these settings fixed:
The Mac and its power mode
The prompt and chat template
The output-token limit
Temperature and sampling settings
Warm-up count
Number of measured requests
The client’s decode speed includes HTTP streaming, so expect it to be slightly lower than the speed each engine prints in its own log. Compare the client numbers with each other, not with the CLI numbers.
The Mirai CLI has no flag to switch speculation off, so this procedure compares complete engine paths rather than isolating DFlash-Weaver.
The llama.cpp runs with and without MTP in the first video are the closest same-engine view of what speculation adds.
Conclusion
Loading a quantized checkpoint is only the first step in local inference.
The actual work is keeping generation useful once every request runs at batch size one on a user’s device.
Uzu handles that problem at several levels.
Quantization reduces memory traffic.
Its Metal kernels target Apple hardware.
DFlash-Weaver proposes longer token sequences, while rollback-free verification checks them without repeatedly restoring recurrent state.
On a base M5, that combination took a 4-bit 9B model from 22 tokens per second with plain llama.cpp to between 92 with Uzu on the same prompt.
Acceptance changes with the workload. Hardware, quantization quality, context length, and thermal state can also alter the result. The correct test uses the application’s prompts and measures quality alongside latency.
If you want to dive deeper, the Uzu repository contains the engine, local server, build instructions, and bindings for Rust, Swift, Python, and TypeScript.
Start with the local server, run the benchmark client from this article, and replace the sample prompt with requests from the actual product.
git clone https://github.com/trymirai/uzu.git(don’t forget to star it ⭐️)
Good day!


















