MTPLX 2.12.2 fixes a regression in 2.12.1 that refused ordinary prompts on Macs with 8, 16 and 32 GB of memory. The memory guard in 2.12.1 charged every model as if it were the 27B and counted the engine's normal memory as a leak, so the model the app picks for those Macs answered many prompts with a 507 "insufficient memory" error. 2.12.2 prices each prompt at what it was measured to use, runs a prompt in smaller chunks before refusing it, and stops charging a small Mac for memory a healthy engine always holds. Macs with 64 GB or more, Flash-Next and the context windows are unchanged.
Highlights
| What you do | 2.12.1 | 2.12.2 |
|---|---|---|
| Send a one-line request to Qwen 3.5 4B on GitHub's 7 GB M1 test machine | Refused with a 507, twice | Served |
| Send an 18-token prompt to Qwen 3.5 4B on an 8 GB Mac with 3 GiB free | Refused with a 507 | Served |
| Send 2,500 and 10,000-token prompts to the 4B on an 8 GB Mac with 3 GiB free | Both refused | Both served |
| Run a five-turn agent session (2,532 to 7,589 tokens) on the 4B, 8 GB Mac, 3 GiB free | Not run | All five turns served, each reusing the previous turn |
| Send 2,600, 6,800 and 7,800-token prompts to Ternary Bonsai 2 27B on a 16 GB Mac | All three refused | All three served |
| Start an agent session on Bonsai 2 27B on a 16 GB Mac | The first turn (2,570 tokens) was refused | Four turns served, 2,580 to 6,467 tokens |
| Send 6,800 and 7,800-token prompts to Qwen 3.8 27B Optimized Speed on a 32 GB Mac | Both refused | Both served |
| Run an agent session on the 27B on a 32 GB Mac | The third turn (about 5,100 tokens) was refused | Four turns served, 2,584 to 6,385 tokens, and also with only 4 GiB free |
| Continue the 99,355-token 27B turn from #499 on a 48 GB Mac | Refused before its prompt was read | Runs in 1,024-token chunks, priced at 34.87 GiB against the 36 GiB limit (tests with the reporter's numbers) |
The first row ran on GitHub's macOS 14 test machine, an M1 with 7 GB, on October 2 (2.12.1) and October 3 (the 2.12.2 code). The other rows come from runs on October 3 on an M5 Max with 128 GB, set up as each smaller Mac: the engine's memory limit for that Mac (8, 12 and 24 GiB), its free-memory floors, a fixed reading of free memory (8 GiB on the 16 and 32 GB rows unless stated), and the model the app picks for that Mac, one model at a time, with 64-token answers. 2.12.1 and 2.12.2 ran the same requests, except the 8 GB agent session. 2.12.0, run on the same requests, served the 16 and 32 GB rows and the 18 and 2,500-token prompts on the 8 GB Mac, and refused the 10,000-token one.
What changed
-
Each prompt is priced at what it really uses. 2.12.1 charged every model the 27B's figure of 3 GiB for each 2,048-token chunk of a prompt, and about 2 GiB even for a one-line prompt. Measured through
mtplx serveon October 2 and 3, a chunk of the 4B uses 0.19 GiB for a 59-token prompt and 1.55 to 1.65 GiB at 2,048 tokens, a chunk of Bonsai 2 27B uses 0.45 GiB for a 98-token prompt, and Bonsai and the 27B use 2.5 to 2.7 GiB at 2,048 tokens. The 4B, the 27B and Bonsai are now charged those amounts plus a margin of 0.09 to 0.25 GiB, and the 9B models by the same rule. Models with routed experts (Qwen 3.6 35B-A3B), Flash-Next and Gemma 4 keep the prices they had in 2.12.1. -
A prompt that does not fit runs in smaller chunks. Before refusing a prompt, the engine now tries 1,024-token chunks and, for prompts up to about 8,000 tokens, 512-token chunks. A 1,024-token chunk uses about 1 GiB less than a 2,048-token one on the 27B models, and the saving holds as the conversation grows (measured to 8,000 tokens on the 27B models and 49,000 on the 4B). A 512-token chunk uses less again on short prompts, but more as the conversation grows, so the engine only picks it where it is cheaper. Smaller chunks are a little slower: in single runs, not a controlled speed test, a 7,000-token prompt on Bonsai took 9.5 s in 2,048-token chunks, 9.6 s in 1,024 and 9.7 s in 512. The engine uses a smaller chunk only when the prompt would not fit otherwise and the smaller chunk saves at least 128 MiB.
-
Memory the engine holds outside the GPU allocator is charged only past 4 GiB on Macs under 64 GB. 2.12.1 charged everything past a sixteenth of RAM, which is 1 GiB on a 16 GB Mac and 2 GiB on a 32 GB Mac. A healthy engine holds 2.1 to 2.9 GiB there with Bonsai and 2.3 to 4.0 GiB with the 27B, so normal use was counted as a leak. Memory past 4 GiB is still charged, so a leak like #546 is still caught. Macs with 64 GB or more keep a sixteenth of RAM, up to 8 GiB.
-
Dependencies. urllib3 is now 2.8.0 (#582).
Validated
| Check | Result |
|---|---|
| Prompt memory of the 4B, Bonsai 2 27B and the 27B, in chunks of 512, 1,024 and 2,048 tokens, prompts from 60 to 49,468 tokens | 43 requests through mtplx serve, the source of every price above |
| Before and after on 8, 16 and 32 GB settings, as in the table | 2.12.1, 2.12.0 and 2.12.2 on the same requests |
| GitHub's 7 GB M1 test machine: download the 4B, start the server, answer a one-line prompt | Passed on the 2.12.2 code; failed with a 507 on 2.12.1 |
| Test suites on the release commit | pytest 11,284 passed (13 new tests for these cases), Swift 1,104 passed, M1 to M4 rehearsal 763 passed, 0 failures |
| The release wheel, installed fresh, on the 8, 16 and 32 GB settings | Same results as the table above |
Known issues
- No 8, 16 or 32 GB Mac has run this release with a full desktop. Apart from GitHub's 7 GB M1 test machine, which runs nothing else, the rows above come from an M5 Max set up as each of those Macs. The limits, floors and prices are the same arithmetic a real Mac of that size uses, but the free memory on a real Mac moves while other apps run.
- With little memory free, longer prompts are still refused. The engine keeps at least 1 GiB free for macOS, as 2.12.0 and 2.12.1 did. With 2 GiB free on an 8 GB Mac, a 2,500-token prompt was refused 0.2 GiB short. With 3 GiB free on a 16 GB Mac, prompts up to 5,100 tokens were served and 6,300 and 6,800-token prompts were refused 0.1 to 0.2 GiB short. Closing other apps frees the memory these need.
- The known issues of 2.12.1 still apply, except two: the 8 GB pricing entry, which this release fixes, and the #499 refusal on 48 GB Macs, which now runs in smaller chunks in the tests. See the 2.12.1 release notes.
Upgrading
Run brew upgrade mtplx or pip install -U mtplx, or use the app's update check. You don't need to change any settings.
Contributors
- urllib3 2.8.0, by @dependabot in #582
Full changelog: https://github.com/youssofal/MTPLX/compare/v2.12.1...v2.12.2
1d7df582c6ca9cf9894fb3642006042ad7f3e5d1462157c13b8e9de736a4ff20. Every published speed
number with its conditions is on the benchmarks page; the archive of every version is
on the releases page.