Skip to content
samanb
Go back

M6 Mac mini and M5 Mac Studio: How Much Do You Need for Local AI?

The M6 chip. Image: Apple Newsroom.

The M6, Apple’s first 2nm chip. Image: Apple.

Buy by memory, not by chip name. Apple put the M5 Mac Studio and M6 Mac mini up for pre-order this week. Neural Engines and “Neural Accelerators” sound impressive, but they do not decide whether a large model loads. Two numbers do: how much unified memory you get, and how fast that memory can move data.

I simulated all 13 configurations through llmfit, an open-source model-to-hardware checker built by Alex Jones, a good friend and former colleague from the London software scene, against the current open-weight lineup as of 28 August: Kimi K3, Qwen3.8-Max, DeepSeek V4, LongCat 2.0, GLM-5.2, Gemma 4, gpt-oss and the rest.

The pattern came out cleaner than I expected:

Four terms, once

What you want → what to buy

Find the row that matches your actual workload. Want the full receipts? Jump to the fit matrix.

What you wantMinimum RAMMachine to look at
Private chat, documents, light coding help16–32GBMac mini M6 (32GB is the comfortable floor)
Serious daily local models (gpt-oss-120b, DeepSeek V4-Flash class)64–128GBMac Studio M5 Max, or M5 Pro mini at 64GB if you only need that model
400B class and mid-pack coding (GLM-5.2, MiniMax M3). Most daily prompts local; subscription downgraded, not cancelled256GBMac Studio M5 Ultra
Every open model that fits on a desk, often at Q8. Still hybrid: Claude or Codex for the hard work. Biggest open downloads (Kimi K3, Qwen3.8-Max, LongCat) stay off-desk too512GBMac Studio M5 Ultra (late October)

Pre-orders are open now, first machines land 22 September, and the 512GB tier arrives in late October.


The fit matrix

In a hurry? The quick pick above is enough. This matrix is the receipt: every model worth running locally, against every machine Apple sells. Find your model, read across the row, and buy the cheapest column with a green tick.

Fit is that quant’s memory footprint against each machine’s usable memory: Perfect under 60%, Good under 85%, Marginal under 98%, Too Tight beyond. A model never rates worse on a bigger machine. Each model is scored on one representative quant, the best from a trusted publisher that runs on the smallest machine it can.

Model fit and estimated tok/s across Apple configs
Model Mac mini Mac Studio
M6 M5 Pro M5 Max M5 Ultra
16GB 24GB 32GB 24GB 48GB 64GB 36GB 48GB 64GB 128GB 96GB 256GB 512GB
giant open
Kimi K3 2.8T
Qwen3.8-Max 2.4T
DeepSeek V4-Pro 1.6T 12
LongCat 2.0 1.6T
Ling 2.6 1T 2.4
GLM-5.2 743B 15 15
Nemotron 3 Ultra 550B 11 11
mid
Llama 4 Maverick 400B 34 34
Qwen3.5 397B-A17B 18 34 34
DeepSeek V4-Flash 284B 23 45 45 45
MiniMax M3 427B 30 59 59
Kimi K2.6 1.1T
Nemotron 3 Super 120B 4.2 7.5 7.5 11 15 15 15 30 30 30
gpt-oss-120b 117B 37 74 74 144 144 144
Llama 4 Scout 109B 8.8 8.8 13 18 18 18 34 34 34
small
Gemma 4 31B 11 12 12 22 22 22 33 44 44 44 86 86 86
Gemma 4 26B-A4B 19 21 21 37 37 37 56 75 75 75 146 146 146
Qwen3.8 27B 7.0 7.7 7.7 14 14 14 21 28 28 28 55 55 55
ERNIE 4.5 21B-A3B 20 20 35 35 35 53 71 71 71 138 138 138
Granite 4.1 30B 12 13 13 23 23 23 35 47 47 47 92 92 92
Nemotron 3 Nano 30B 23 26 26 47 47 47 70 93 93 93 182 182 182
Ministral 3 14B 8.1 8.9 8.9 16 16 16 24 32 32 32 63 63 63
Gemma 4 12B 14 16 16 28 28 28 42 56 56 56 110 110 110
gpt-oss-20b 21B 16 17 17 31 31 31 47 62 62 62 122 122 122

Legend: = fits and usable · = compromised · = won't fit · the small number is the estimated tok/s.

Reading the table. Green means the model loads; the small numbers are decode tok/s from llmfit, so compare columns rather than treating any one cell as a lab result. The 512GB Ultra is the desk ceiling for open weights that fit, not a Claude or Codex replacement: DeepSeek V4-Pro and Ling 2.6 only scrape in at demo speeds there. GLM-5.2 runs from 256GB up. The largest open downloads, Kimi K3, Qwen3.8-Max, LongCat 2.0 and Kimi K2.6, are red at every price Apple charges (K3 alone is 2.8T): too big for any desk Mac. Ultra 96GB is the odd one out, more bandwidth than Max 128GB but less memory. A tick marks the quant that fits, 2-bit to 4-bit at the tight end, Q8 at the top: quality is a trade there.

The table shows decode only. Prompt processing, the time to first token, is compute-bound, and that is where Apple’s claimed 4× jump lives. On the last generation, extra GPU cores bought roughly 22% more prompt processing but 6% more decode (llama.cpp benchmarks). Feed it long documents and agent tool output, and the £1,300 step to the 80-core Ultra earns itself in ways the tok/s column cannot see.


What you actually get for your money

You have the RAM tier from the matrix. The next question is what that spend actually buys: which models run comfortably, how fast, and whether a cloud subscription still earns its keep.

Apple’s UK store lists the Mac Studio M5 Max from £2,499 (36GB) and £3,099 (48GB), the M5 Ultra from £5,499 (96GB), the Mac mini M6 from £899, and the M5 Pro mini from £1,699. The 256GB Ultra is already configurable there; the 96GB to 256GB memory upgrade adds £4,000 on Apple’s price ladder, which puts the entry 256GB build at £9,499 before CPU or storage steps. The 512GB option is listed for late October with no published price yet. Apple dropped the 512GB M3 Ultra during this year’s memory shortage (Cult of Mac), so watch the order page when that tier opens.

How good is the code it writes? I measure that with SWE-bench Pro: models get real coding tasks from 41 software projects, and the score is the percentage they finish correctly. It replaced an older benchmark that stopped meaning much once the test questions leaked into training data. On the August 2026 vendor board, the hosted tip readers usually mean, Claude Fable 5 and Mythos Preview, sit around 80%, with GPT-5.6 Sol at 64.6% and Grok 4.5 at 64.7%. Open-weight Qwen3.8 Max posts 67.7% but is too big for these Macs. The best open weights you can actually load sit in the mid-pack: GLM-5.2 at 62.1%, MiniMax M3 at 59.0%, DeepSeek V4-Pro at 55.4%. Local clears last season’s GPT-5.5 (58.6%) and sits with GPT-5.6 Luna / Claude Sonnet 5; it does not match Fable 5 or Codex-class work. That is why the hybrid pattern below still holds.

What you get at each budget: machine, top speed, best daily drivers and the honest ceiling
BudgetMachineTop speedBest daily drivers
Entry Mac mini 16GB, from £899 up to 16 tok/s gpt-oss-20b (16), Gemma 4 12B (14)

Chat, document Q&A over your own files (RAG: the model looks things up in your documents instead of guessing), and light coding help. Private and offline. Your subscription still earns its keep for real software work.

Compact Mac mini 32GB, about £1,200 (est) up to 21 tok/s Gemma 4 26B-A4B (21), gpt-oss-20b (17)

The same jobs with more headroom, and one genuine surprise: Gemma 4 26B scores 82.3% on GPQA Diamond, graduate-level science, from a box this size. Still not an agentic coder (a setup where the model plans, calls tools, and loops on a task without you steering every step), so keep the subscription.

Sweet spot Studio Max 64 to 128GB, £2,499 and up (£6,899 at 4TB) up to 74 tok/s gpt-oss-120b (74), DeepSeek V4-Flash (23)

Speeds shown are the 128GB figures; a 64GB Max sits between this tier and the mini. Fast local coding help on an Apache 2.0 model that runs fully in memory. Useful for agent loops, not a Claude replacement: see Can I cancel my Claude subscription?. This is where the subscription starts shrinking: the daily 80% of prompts stop costing a monthly fee.

400B class Studio Ultra 256GB, £9,499 up to 59 tok/s MiniMax M3 (59), Qwen3.5 397B (34), GLM-5.2 (15)

GLM-5.2 at 62.1 SWE-bench Pro sits with GPT-5.6 Luna (62.7) and Claude Sonnet 5 (63.2), ahead of GPT-5.5 at 58.6 (vendor board). MiniMax M3 posts 59.0 on the same board. Neither is Fable 5 (80). No API, no per-token bill for that mid-pack work.

Desk ceiling Studio Ultra 512GB, £15,000 to £16,500 (est, late October) up to 15 tok/s GLM-5.2 at Q8 (15), Nemotron 3 Ultra (11), DeepSeek V4-Pro (12)

Everything in the open-weight world that fits on a desk, often at Q8. Speeds above are the slow giants this tier unlocks or cleans up; MiniMax still runs at 59 here. The biggest open downloads (Kimi K3, Qwen3.8-Max, LongCat) still will not load. The subscription stays: this is not a Claude or Codex replacement.

Two patterns run through these numbers. Across chips, money buys a bigger model and more speed: the same Gemma 4 26B-A4B decodes at 19 tok/s on the 16GB mini and 146 on the 512GB Ultra, because memory bandwidth climbs from 153GB/s to 1.2TB/s. Within one chip, memory buys a bigger model, and speed follows active parameters instead: the Entry tier’s gpt-oss-20b outruns the Compact tier’s Qwen3.8 27B on the same M6, 16 tok/s against 8, because it is the smaller model. Memory buys model size; bandwidth and active parameters buy speed.

One row worth a second look before you jump to the Studio: the M5 Pro mini at 64GB runs gpt-oss-120b at 37 tok/s for £1,899, roughly half the Max’s entry price. If the sweet-spot model is the draw and the Studio’s extra memory is not, the Pro mini is the bargain in this lineup.

SWE-bench Pro scores above are full-precision figures; the quants that fit smaller machines trade some of that away. GLM-5.2’s 62.1 holds up well on the 512GB Ultra. Use the scores as an upper bound for what that model class can do.


Can I cancel my Claude subscription?

Not on 128GB, and the payback maths is slower than the marketing suggests on anything bigger.

Local does not replace the subscription, it shrinks it. At 128GB and up, the daily 80% of prompts, chat, RAG, boilerplate, refactors, document processing, moves onto hardware you already own at zero marginal cost. What is left is the hard 20% where you still want Claude, Codex or whatever hosted tip you already trust. You drop from a Max plan to a Pro plan, or from Pro to occasional API top-ups, not to zero.

The reverse also applies. A 512GB Ultra running GLM-5.2 at 15 tok/s is a strong local mid-pack setup (Sonnet 5 / Luna territory on SWE-bench Pro), and it is still slower per token than a cheap cloud API. You are paying £15,000 or more upfront for privacy, offline access and no metering, not for speed. Below 128GB the compact models are impressive for their size and they are not Claude.

I run this workload on a 128GB Strix Halo desktop. gpt-oss-120b was too slow and not accurate enough for my use: time to first token is noticeable on long prompts, and tok/s sits well behind a cloud subscription. The best model I have tried on that box is Qwen, and even it is a step down from Claude or Codex. Agentic coding is not one agent either: long contexts, tool outputs and parallel sub-agents each want their own context window, and a machine that comfortably holds one long agentic context does not comfortably hold five. The agentic engineering terms for sub-agents, context windows and the loop are why 128GB looks fine on paper and cramped on a desk. Speed compounds the problem: 74 tok/s is one agent’s speed, not five agents sharing the machine. My verdict for this camp: 256GB is the floor. If you plan to replace Claude or Codex with 128GB, you will be disappointed. I own a 128GB machine.

Where that leaves the five types of buyer:

DemographicMachineVerdict
The tinkerermini 16 to 32GBChat, RAG, docs. Keeps the subscription.
The privacy-focused professionalmini 32GB to Max 64GBLocal for sensitive repos, subscription for the rest.
The daily-driver optimiserUltra 256GBMost coding work local. Subscription downgraded, not cancelled.
The agentic power userUltra 256GB minimum, 512GB preferredHybrid pattern, modest real savings.
The local-first zealotUltra 512GBDesk ceiling at Q8. Still hybrid: Claude or Codex for the hard 20%. Biggest open downloads still will not fit.

The payback maths, since somebody will run it. Anthropic’s pricing page puts the top subscription at about £200 a month. Local does not cancel that bill; the honest cases are a downgrade, or an aggressive cut to a cheap plan. Against Studio Ultra prices, subscription savings alone are a slow payback:

Subscription payback against Mac Studio Ultra hardware: scenario, hardware price, sub change, annual save, rough payback years
Scenario Hardware Sub change Annual save Rough payback
Full cut to £20 Ultra 256GB, £9,499 £200 → £20 £2,160 ~4.4 years
Downgrade Max → Pro-ish Ultra 256GB, £9,499 £200 → £90 £1,320 ~7.2 years
Downgrade Max → Pro-ish Ultra 512GB, ~£15,500 £200 → £90 £1,320 ~12 years

Payback = hardware ÷ annual save. Figures use Anthropic list prices as of writing; mid-point used for the 512GB estimate (£15,000–£16,500). The highlighted row is the realistic hybrid path.

Hardware is rarely bought on subscription savings alone. Privacy, offline use, no rate limits and no metering are the real reasons; the savings are a bonus that eventually arrives.

The £20 plan still earns its keep next to a big local box: Claude and Codex on rate-limited plans run out fast on long tasks and parallel sub-agents. The realistic pattern is hybrid, local carries the bulk and the subscription tops up the hard 20%. Nobody should buy a £15,000 computer to cancel a £200 subscription. Plenty will buy it to own the stack they can run locally.


Specs for the curious

The matrix columns are these machines. Chip name is the marketing line; RAM and bandwidth are the ones that decide fit and decode speed.

Two machines, 13 configurations. The Mac Studio ships M5 Max and M5 Ultra; the mini ships M6 (not M5) plus an M5 Pro upgrade. Apple’s store lists four mini configurations, where models 1 and 2 differ only in storage. One genuine SKU quirk: the 16GB M6 runs 153GB/s against 170GB/s for the 24 and 32GB configs, so the cheapest mini gives up a little bandwidth as well as memory.

Specs come from the Mac Studio and Mac mini pages. Prices are Apple’s official UK figures from the store pages linked above.

M5 Max/Ultra Mac Studio

Mac Studio handling agentic AI workloads. Image: Apple.

ConfigChipCPU/GPURAMBandwidth
M5 Max 36GB baseM5 Max18c / 32c36GB460GB/s
M5 Max 48GBM5 Max18c / 40c48GB614GB/s
M5 Max 64GBM5 Max18c / 40c64GB614GB/s
M5 Max 128GBM5 Max18c / 40c128GB614GB/s
M5 Ultra 96GB baseM5 Ultra30c / 64c96GB1.2TB/s
M5 Ultra 256GBM5 Ultra36c / 80c256GB1.2TB/s
M5 Ultra 512GBM5 Ultra36c / 80c512GB1.2TB/s

M5 Pro/M6 Mac mini

Mac mini running on-device AI agents. Image: Apple Newsroom.

ConfigChipCPU/GPURAMBandwidth
M6 16GB (models 1/2)M612c / 12c16GB153GB/s
M6 24GB (model 3)M612c / 12c24GB170GB/s
M6 32GB (model 3 RAM upgrade)M612c / 12c32GB170GB/s
M5 Pro 24GB (model 4 base)M5 Pro15c / 16c24GB307GB/s
M5 Pro 48GB (model 4)M5 Pro15c / 16c48GB307GB/s
M5 Pro 64GB (model 4, 18c/20c upgrade)M5 Pro18c / 20c64GB307GB/s

How much storage do you actually need?

RAM decides whether a model loads. SSD decides whether you can keep more than one. Sweet-spot weight files already run past 100GB; Ultra giants sit in the 300–450GB range, and downloads need spare room whilst they land.

Comfortable sizes below come from llmfit’s disk_size_gb (three largest fitting models + 100GB OS + download scratch, rounded up to an Apple SKU). Workings: disk_comfort.py and the comfort table.

Comfortable Apple SSD size by buyer tier, from disk need formula
Tier Need Comfortable SSD
Entry / Compact 158–212 GB 256 GB
Pro mini / Max 36–64 GB 232–329 GB 512 GB
Sweet spot (Max 128 GB) 566 GB 1 TB
Ultra 96 GB 440 GB 1 TB
400B class (Ultra 256 GB) 1,073 GB 2 TB
Desk ceiling (Ultra 512 GB) 1,882 GB 4 TB

Entry / Compact: bump to 512 GB if you keep more than three small models on disk.

Apple’s base SKUs undershoot on most tiers: mini often 256GB, Pro mini and Max 512GB, Ultra 1TB. Neither RAM nor SSD is upgradable later, so configure storage when you order. Cold models can sit on a Thunderbolt drive; keep the active set internal.



How I ran the numbers

llmfit scores every model on Hugging Face against your hardware, estimates decode speed from memory bandwidth scaled to each chip’s published figure, and its override flags simulate any machine: llmfit --memory 512G --ram 512G --cpu-cores 36 fit --json. Two judgement calls shaped the tok/s column. Mixture-of-experts models read only their active experts per token (about 5.1B of gpt-oss-120b’s 117B), so they run far faster than their file size suggests; I scored those from active parameters rather than total download. Where llmfit’s own speed model runs conservative, I calibrated against published measurements and kept the lower end. SSD recommendations reuse each fit row’s disk_size_gb through disk_comfort.py: top three fitting models + 100GB OS + download scratch, round up to an Apple SKU, bump one tier if that drive would be more than 85% full. The thirteen runs behind this post ship as raw JSON in the analysis repo.



Reproduce it yourself

If you want to check my work, install llmfit and run the same flags.

macOS (Homebrew or uv/pipx):

brew install uv           # or: brew install pipx
uv tool install llmfit    # or: pipx install llmfit

Linux (uv or pipx):

uv tool install llmfit    # or: pipx install llmfit

Windows (WSL or native Python):

pipx install llmfit

Then simulate any config:

llmfit --memory 512G --ram 512G --cpu-cores 36 fit --json

Drop the --json and llmfit opens an interactive TUI instead. Press S to change the simulated hardware live and watch every verdict re-rank.

llmfit --memory 128G --ram 128G --cpu-cores 18

It is a handy way to sanity-check any config before committing to it.

The thirteen simulation runs behind this post are in the mac-studio-m5-analysis repo as raw JSON, along with the per-config docs and the command list.

If your model class is already covered at 128GB, the Max is the sensible buy. If you need the 400B class or want the mid-tier giants at comfortable speeds, it is the Ultra.


Share this post

Comments


Next Post Your AI Coding Assistant Is Burning Through Your Budget. Here's Why.