Native MTP · Qwen 3.8 Flash Next · Apple Silicon · v2.11.3

Run local LLMs
twice as fast.

MTPLX is the fastest way to run Qwen 3.8 on a Mac. A native app for local AI on Apple Silicon: chat, coding agents, and an OpenAI and Anthropic compatible local server, with native MTP speculative decoding and output that stays exact at any temperature. On an M5 Max, at each model's own sampler: Qwen 3.8 Flash Next at 125 tok/s on an OpenCode request and 79 tok/s at 9k tokens of context, Qwen 3.8 27B up to 87.6 tok/s, and 81.74 tok/s on Qwen 3.6 27B with the raw logs published.

macOS 14+ · Apple Silicon · Free & open source
Same prompt · Recorded in real time

Twice as fast. Still exact.

Twice the speed at no quality loss, at any temperature, and up to 3x on the 8-bit Quality pack. The fast path is checked against the plain path by token id on every release. The 27B record: 81.74 tok/s on Qwen 3.6 27B at 2.69x plain decode, M5 Max, raw logs on Benchmarks. Qwen 3.8 Flash Next: 125.8 tok/s on an OpenCode request, 79.3 tok/s at 9k context, 61.8 at 109k, on MTPLX 2.11.3. How it works →

MTP off
MTP on
Qwen 3.8 on a Mac · MTPLX 2.11.3

The fastest way to run Qwen 3.8 Flash Next on a Mac.

Qwen 3.8 Flash Next is the 125B mixture of experts. MTPLX shipped the first Apple Silicon backend for it and runs its own MTP head as an exact speculative decoder. Measured on a MacBook Pro M5 Max with 128 GB, fans verified, at the model's own sampler.

01 · One-click launch

Your favorite tools, at twice the speed.

One click serves OpenCode, Pi, Hermes, or the web UI from your Mac. OpenAI and Anthropic compatible, so everything plugs in.

02 · Chat

Chat, built in.

Native Swift, fully offline, in twelve languages, dark or light. Watch your models write code and prose at twice the speed.

03 · Forge

Forge fast MTP models.

Paste a Hugging Face link. Forge converts it to MLX and measures the speedup on your Mac.

Get MTPLX

Ready in one download.

DMG · 60 MB · everything included