MTPLX is the fastest way to run Qwen 3.8 on a Mac. A native app for local AI on Apple Silicon: chat, coding agents, and an OpenAI and Anthropic compatible local server, with native MTP speculative decoding and output that stays exact at any temperature. On an M5 Max, at each model's own sampler: Qwen 3.8 Flash Next at 125 tok/s on an OpenCode request and 79 tok/s at 9k tokens of context, Qwen 3.8 27B up to 87.6 tok/s, and 81.74 tok/s on Qwen 3.6 27B with the raw logs published.
Twice the speed at no quality loss, at any temperature, and up to 3x on the 8-bit Quality pack. The fast path is checked against the plain path by token id on every release. The 27B record: 81.74 tok/s on Qwen 3.6 27B at 2.69x plain decode, M5 Max, raw logs on Benchmarks. Qwen 3.8 Flash Next: 125.8 tok/s on an OpenCode request, 79.3 tok/s at 9k context, 61.8 at 109k, on MTPLX 2.11.3. How it works →
Qwen 3.8 Flash Next is the 125B mixture of experts. MTPLX shipped the first Apple Silicon backend for it and runs its own MTP head as an exact speculative decoder. Measured on a MacBook Pro M5 Max with 128 GB, fans verified, at the model's own sampler.
One OpenCode request on the Optimized Speed pack: 1,301 tokens generated, 18,539-token prompt with 18,364 tokens served from the cache, MTP depth 3. 16 September 2026.
79.3 tok/s on a 9k-token code prompt, 27 percent above 2.11.2 on the same Mac and runtime. 61.8 tok/s on a 109k-token OpenCode turn, 50.3 at 200k.
87.6 tok/s rewriting a file it just wrote, 65.2 on a fresh coding task at Qwen's official sampling. The 4-bit pack agrees with bf16 on 96.0 percent of top-1 tokens.
mlx-serve's own M5 Max table puts MTPLX exact decoding at 102.0 tok/s against its own 92.5. Every comparison with a version, a machine and a date on each number.
One click serves OpenCode, Pi, Hermes, or the web UI from your Mac. OpenAI and Anthropic compatible, so everything plugs in.
Native Swift, fully offline, in twelve languages, dark or light. Watch your models write code and prose at twice the speed.
Paste a Hugging Face link. Forge converts it to MLX and measures the speedup on your Mac.