Post

Konversation

This inference engine runs LLMs 4x faster than llama.cpp: (Qwen3.5-9B at 92 tokens/s on 16GB Mac) I ran Qwen3.5 9B on a base M5 MacBook Pro using three local inference engines, one at a time. → llama.cpp reached 22.0 tokens/s → MLX reached 25.1 tokens/s → Uzu reached 92.1 tokens/s The video shows all three runs side by side. The first two results do not mean llama.cpp or MLX are poorly optimized. Both are approaching a physical limit of ordinary local decoding. Generating each new token normally requires another pass through the model. That pass streams roughly 5.2 GB of weights through memory. A base M5 provides about 153 GB/s of memory bandwidth. At one token per pass, the theoretical ceiling is roughly 153 ÷ 5.2, or 29 tokens/s, however fast the GPU computes. Uzu gets past that ceiling by optimizing speculative decoding for Apple Silicon. Speculative decoding uses a smaller drafting system to propose several future tokens. The full 9B model then checks those proposals together instead of generating every token through a separate pass. Uzu pushes this further in two ways. → It drafts up to 16 positions and uses Weaver to turn independent guesses into coherent sequences. → It verifies several possible sequences as a tree, then commits state only for the path the full model accepts. The 9B model remains in control of the final output. Uzu simply gives it more useful work to approve in each pass. In my run, every pass of the 9B model produced 7.5 output tokens on average. Ordinary decoding produces exactly 1. That is how Uzu reached 92.1 tokens/s on hardware whose one-token-per-pass ceiling is roughly 29. The rest of the engine is designed around this path too. Uzu ships its own quantized checkpoints, and on M5 chips its Metal kernels run verification through the GPU’s accelerated int8 matrix path. It is an open-source inference engine from Mirai, with bindings for Rust, Swift, Python, and TypeScript. The repository is here: github.com/trymirai/uzu (don't forget to star 🌟) I wrote the full breakdown of why local inference is memory-bound and how Uzu gets past that limit, with reproducible results and code. The article is quoted below.
0:00 / 0:26
Zitat
Akshay 🚀
@akshay_pachaar
Cover-Bild für Artikel
Karpathy’s Trick for Faster Local LLMs, finally has a proper solution
In 2023, Karpathy explained why local LLMs are slow. When a model generates one stream of text on your own computer, the chip spends most of its time waiting for weights to arrive from memory while...
N2k1
Deine Antwort posten

Übersetzt (Originalsprache Englisch)
Die Bandbreitenrechnung stimmt, 153 geteilt durch 5,2 ergibt etwa 29. Aber das vergleicht spekulative Dekodierung mit einfacher Dekodierung, und llama.cpp unterstützt ebenfalls spekulative Dekodierung. Der Gewinn hängt davon ab, wie oft vorgeschlagene Tokens akzeptiert werden,
Übersetzt (Originalsprache Englisch)
92 Tokens/s ist das Ergebnis dieses Prompts, kein universeller Wert. Uzu lag bei wiederholten Durchläufen dieses Befehls zwischen 92 und 117 Tokens/s, aber die Akzeptanz variiert je nach Arbeitslast.
One thing before people swap engines: uzu runs its own model format, so your existing GGUFs don't carry over. You convert with their lalamo tool or pull their prebuilt checkpoints. Nice part is it ships an OpenAI-compatible server, so pointing an agent at it is easy.

Aktuelle Trends

Was gibt’s Neues?

Trend in Frankreich
Le 17
Trend in Frankreich
Pérou
Trends
Doppelmoral
Trend in Frankreich
#BlocusUniversités