Your model runs and the words crawl out. Before buying anything, it is worth knowing that speed here obeys one law: producing a single token means streaming essentially the whole model through the processor, so your ceiling is memory bandwidth divided by model size. A machine with twice the bandwidth is roughly twice as fast on the same model, and a model half the size is roughly twice as fast on the same machine. Every lever below moves one of those two numbers.
You can put a figure on your own ceiling in about a minute. The hardware index lists the memory bandwidth of each machine, and every model page lists the size of the file at each compression level. Bandwidth divided by file size is an upper bound you will not beat, and real output lands some way under it, but it tells you quickly whether you are chasing a small improvement or the wrong machine entirely.
The usual culprit is more specific than that, and it is worth ruling out first because it is invisible. When a model very nearly fits in graphics memory, the tools quietly leave the remainder in ordinary memory and carry on. Speed collapses, often to a tenth, and everything still works: no error, no warning, just a model that has become unusable while looking perfectly healthy. The answer is to make it fit. A smaller compression of the same model that lives entirely in graphics memory beats a larger one that spills, by far more than the difference in quality between them costs you. Check your machine against the model before you conclude anything about either.
Context is the lever people forget. The working memory for a conversation grows turn by turn, and past a point it eats the headroom the model needed and slows every step after. A chat that has got slower the longer it ran is doing exactly what it appears to be doing, so start a fresh one, or trim what the tool re-sends each turn.
Stepping down a model size is the lever people resist, and it is the one that most often works. Quality per parameter has improved quickly enough that this year's mid-sized models do work last year's large ones did, at several times the speed, and the ratings on each model page will tell you what a step down actually costs you before you assume you need the big one. One family bends the law outright: a mixture-of-experts design stores a great many parameters and puts only a few billion of them to work on each token, so it answers at closer to a small model's pace while still asking for a large model's memory. Where speed is your constraint and memory is not, that trade is a good one.
If the model fits, the context is short, the compression is sensible and it is still too slow, then the bandwidth law is what you are hearing. Bandwidth is the specification to shop on for this work, and the hardware index leads with the figure that spec sheets bury, per platform and per machine.
Where next: Check your machine · VRAM and what fits · What compression does to a model