Two numbers decide what a Mac does with local models, and you commit to both at checkout. Unified memory decides what fits, because a model either loads or it does not, and on a Mac the whole machine's memory is available to it, which is the quiet superpower of the platform. Memory bandwidth decides how fast the words come out once it has loaded. Everything else on the spec sheet sits downstream of those two.
The chip sets both, and the enclosure mostly decides which chips you are allowed. Take the M5 chips, the newest Apple generation in our hardware index: a base chip tops out at 32 GB, a Pro doubles that to 64, and a Max doubles it again to 128, roughly doubling the memory bandwidth at each step. Past a Max there is only the Ultra in a Mac Studio, and the Ultra we hold is the M3 of March 2025, two generations behind the laptops and the one part in the index that reaches 512 GB. Every M5 machine the index names so far is a MacBook Pro. Look back a generation for the rest of the line as we hold it: the Air and the mini carry a base chip, the mini also comes with a Pro, and the Ultra lives in the Studio.
Read those ceilings as slightly smaller than they look. macOS needs room, and the conversation needs room on top of the model file, so a model whose file fits your memory on paper does not necessarily fit in practice. That is why the subtraction people do in their heads, memory minus file size, keeps giving the wrong answer, and why the fit check is worth running before you pay for anything. The sharpest illustration sits at the top of the range: the largest open-weights models still spill into slower memory on the biggest Studio we track, at the compression most people use.
The Air is the machine you may already own, and a fine way to find out whether you care about any of this. It runs small models well, it stops at 32 GB, and it has no fan, so long generations slow down as the machine warms. Good for finding out; a stretch as a daily tool.
A MacBook Pro is the laptop answer once local models are part of your working day, and the choice inside the line matters more than the badge on it. A Pro chip with the memory filled will hold models that a Max chip with less of it turns away, and fit comes before speed every time.
The Mac mini is the most capability for the money, with one trick that changes the question entirely: it does not have to be your main computer. A mini running Ollama serves models to every other machine in the house, so your laptop points its tools at the mini's address and stays cool and quiet. If you were pricing a bigger laptop purely for inference, a modest laptop plus a dedicated mini often beats it.
The Mac Studio exists here for memory, and the honest version of that sentence is narrower than the marketing. The big Ultra puts the mid-sized and large open models comfortably in reach while the very largest still spill over, and its own page opens by saying how many of the models we track it holds in memory. If you cannot name the big model you intend to run, you do not need one yet.
The sizing question worth asking is what job the model is doing, because the three common jobs tolerate slowness very differently:
- Writing and chat — replacing a chat subscription for drafting, rewriting, thinking out loud — is the forgiving one. Text needs to arrive at reading speed and no faster, and mid-sized models are genuinely good at prose now. This works on the widest range of machines, including a base chip with modest memory.
- Coding in an editor — hooking a local model up to Cursor, Zed or something similar — is the demanding one, in a way that surprises people. It is not only the words coming out: an editor stuffs your files into every request, and reading that long prompt in is where Apple silicon is weakest relative to a discrete graphics card. The wait before the first word grows with your codebase. Capable local coding models want the larger-memory Pro and Max configurations, and even then expect a pause a hosted service would not have.
- Agentic workflows — chains of steps, tool calls, long contexts — compound both problems: every step re-reads a growing context, small models drop instructions that large ones follow, and each pause multiplies across the chain. This is the use case that justifies the expensive end of the range, and the one where being honest with yourself about what you will really run saves the most money.
Which leaves the practical order. Pick the models you want to run, then let the fit verdicts pick the memory: check any machine against the catalogue, or open a chip's page in the hardware index to see what it holds at the standard compression level and roughly how fast it goes. The four picks below are four rungs of that ladder, each one a link to its own verdict.
Where next: Check what a machine runs · Run your first model on it