MTPLX against Ollama, LM Studio, llama.cpp, mlx-lm and oMLX: same-machine numbers where MTPLX ran both engines, third-party numbers credited to their authors. MTPLX is a free, open source Mac app and CLI that runs local LLMs on Apple Silicon with native MTP speculative decoding. Created by Youssof Altoukhi, who brought native MTP to the Mac in April 2026.
Three published Ollama runs of Qwen 3.8 27B Q4_K_M on other people's Macs beside MTPLX's own Qwen 3.8 27B rows. Different machines and different prompts, so the table shows scale.
Both engines on one M5 Max over a 52,740-token Qwen 3.8 27B answer at the official Qwen 3.8 sampling, 15 Aug 2026: 17.40 tok/s against 32.4. A separate draft model beside a native MTP head.
Two third-party pairs, each on one machine: a MacBook in May 2026 and Mirai Labs' M5 Max in September 2026. Plus the dated MTP history of both engines.
Native MTP beside plain MLX decode of the same model: the v0.1.5 README pair, the 2.9.0 pack-table multipliers, Mirai Labs' board row, and the status of mlx-lm's MTP pull request.
oMLX 0.5.7 and MTPLX 2.7.0 on one M5 Max with the same prompt and sampling, a third-party temperature-0 row, and the dates of oMLX's Lightning MTP running on MTPLX kernels.