MTPLX 2.12.2

Released 2026-10-03. A fix for Macs with 8, 16 and 32 GB. 2.12.1 refused ordinary prompts on these Macs with the model the app picks for them, even a one-line prompt on an 8 GB Mac with 3 GiB free. 2.12.2 prices each prompt at what it was measured to use and serves them.

MTPLX 2.12.2 fixes a regression in 2.12.1 that refused ordinary prompts on Macs with 8, 16 and 32 GB of memory. The memory guard in 2.12.1 charged every model as if it were the 27B and counted the engine's normal memory as a leak, so the model the app picks for those Macs answered many prompts with a 507 "insufficient memory" error. 2.12.2 prices each prompt at what it was measured to use, runs a prompt in smaller chunks before refusing it, and stops charging a small Mac for memory a healthy engine always holds. Macs with 64 GB or more, Flash-Next and the context windows are unchanged.

Highlights

What you do 2.12.1 2.12.2
Send a one-line request to Qwen 3.5 4B on GitHub's 7 GB M1 test machine Refused with a 507, twice Served
Send an 18-token prompt to Qwen 3.5 4B on an 8 GB Mac with 3 GiB free Refused with a 507 Served
Send 2,500 and 10,000-token prompts to the 4B on an 8 GB Mac with 3 GiB free Both refused Both served
Run a five-turn agent session (2,532 to 7,589 tokens) on the 4B, 8 GB Mac, 3 GiB free Not run All five turns served, each reusing the previous turn
Send 2,600, 6,800 and 7,800-token prompts to Ternary Bonsai 2 27B on a 16 GB Mac All three refused All three served
Start an agent session on Bonsai 2 27B on a 16 GB Mac The first turn (2,570 tokens) was refused Four turns served, 2,580 to 6,467 tokens
Send 6,800 and 7,800-token prompts to Qwen 3.8 27B Optimized Speed on a 32 GB Mac Both refused Both served
Run an agent session on the 27B on a 32 GB Mac The third turn (about 5,100 tokens) was refused Four turns served, 2,584 to 6,385 tokens, and also with only 4 GiB free
Continue the 99,355-token 27B turn from #499 on a 48 GB Mac Refused before its prompt was read Runs in 1,024-token chunks, priced at 34.87 GiB against the 36 GiB limit (tests with the reporter's numbers)

The first row ran on GitHub's macOS 14 test machine, an M1 with 7 GB, on October 2 (2.12.1) and October 3 (the 2.12.2 code). The other rows come from runs on October 3 on an M5 Max with 128 GB, set up as each smaller Mac: the engine's memory limit for that Mac (8, 12 and 24 GiB), its free-memory floors, a fixed reading of free memory (8 GiB on the 16 and 32 GB rows unless stated), and the model the app picks for that Mac, one model at a time, with 64-token answers. 2.12.1 and 2.12.2 ran the same requests, except the 8 GB agent session. 2.12.0, run on the same requests, served the 16 and 32 GB rows and the 18 and 2,500-token prompts on the 8 GB Mac, and refused the 10,000-token one.

What changed

  • Each prompt is priced at what it really uses. 2.12.1 charged every model the 27B's figure of 3 GiB for each 2,048-token chunk of a prompt, and about 2 GiB even for a one-line prompt. Measured through mtplx serve on October 2 and 3, a chunk of the 4B uses 0.19 GiB for a 59-token prompt and 1.55 to 1.65 GiB at 2,048 tokens, a chunk of Bonsai 2 27B uses 0.45 GiB for a 98-token prompt, and Bonsai and the 27B use 2.5 to 2.7 GiB at 2,048 tokens. The 4B, the 27B and Bonsai are now charged those amounts plus a margin of 0.09 to 0.25 GiB, and the 9B models by the same rule. Models with routed experts (Qwen 3.6 35B-A3B), Flash-Next and Gemma 4 keep the prices they had in 2.12.1.

  • A prompt that does not fit runs in smaller chunks. Before refusing a prompt, the engine now tries 1,024-token chunks and, for prompts up to about 8,000 tokens, 512-token chunks. A 1,024-token chunk uses about 1 GiB less than a 2,048-token one on the 27B models, and the saving holds as the conversation grows (measured to 8,000 tokens on the 27B models and 49,000 on the 4B). A 512-token chunk uses less again on short prompts, but more as the conversation grows, so the engine only picks it where it is cheaper. Smaller chunks are a little slower: in single runs, not a controlled speed test, a 7,000-token prompt on Bonsai took 9.5 s in 2,048-token chunks, 9.6 s in 1,024 and 9.7 s in 512. The engine uses a smaller chunk only when the prompt would not fit otherwise and the smaller chunk saves at least 128 MiB.

  • Memory the engine holds outside the GPU allocator is charged only past 4 GiB on Macs under 64 GB. 2.12.1 charged everything past a sixteenth of RAM, which is 1 GiB on a 16 GB Mac and 2 GiB on a 32 GB Mac. A healthy engine holds 2.1 to 2.9 GiB there with Bonsai and 2.3 to 4.0 GiB with the 27B, so normal use was counted as a leak. Memory past 4 GiB is still charged, so a leak like #546 is still caught. Macs with 64 GB or more keep a sixteenth of RAM, up to 8 GiB.

  • Dependencies. urllib3 is now 2.8.0 (#582).

Validated

Check Result
Prompt memory of the 4B, Bonsai 2 27B and the 27B, in chunks of 512, 1,024 and 2,048 tokens, prompts from 60 to 49,468 tokens 43 requests through mtplx serve, the source of every price above
Before and after on 8, 16 and 32 GB settings, as in the table 2.12.1, 2.12.0 and 2.12.2 on the same requests
GitHub's 7 GB M1 test machine: download the 4B, start the server, answer a one-line prompt Passed on the 2.12.2 code; failed with a 507 on 2.12.1
Test suites on the release commit pytest 11,284 passed (13 new tests for these cases), Swift 1,104 passed, M1 to M4 rehearsal 763 passed, 0 failures
The release wheel, installed fresh, on the 8, 16 and 32 GB settings Same results as the table above

Known issues

  • No 8, 16 or 32 GB Mac has run this release with a full desktop. Apart from GitHub's 7 GB M1 test machine, which runs nothing else, the rows above come from an M5 Max set up as each of those Macs. The limits, floors and prices are the same arithmetic a real Mac of that size uses, but the free memory on a real Mac moves while other apps run.
  • With little memory free, longer prompts are still refused. The engine keeps at least 1 GiB free for macOS, as 2.12.0 and 2.12.1 did. With 2 GiB free on an 8 GB Mac, a 2,500-token prompt was refused 0.2 GiB short. With 3 GiB free on a 16 GB Mac, prompts up to 5,100 tokens were served and 6,300 and 6,800-token prompts were refused 0.1 to 0.2 GiB short. Closing other apps frees the memory these need.
  • The known issues of 2.12.1 still apply, except two: the 8 GB pricing entry, which this release fixes, and the #499 refusal on 48 GB Macs, which now runs in smaller chunks in the tests. See the 2.12.1 release notes.

Upgrading

Run brew upgrade mtplx or pip install -U mtplx, or use the app's update check. You don't need to change any settings.

Contributors

  • urllib3 2.8.0, by @dependabot in #582

Full changelog: https://github.com/youssofal/MTPLX/compare/v2.12.1...v2.12.2

Get it. Download the current DMG, or brew upgrade mtplx / pip install -U mtplx. The SHA-256 of MTPLX-2.12.2.dmg is 1d7df582c6ca9cf9894fb3642006042ad7f3e5d1462157c13b8e9de736a4ff20. Every published speed number with its conditions is on the benchmarks page; the archive of every version is on the releases page.