How to · Getting started

Running your first model locally, with Ollama or LM Studio

Fifteen minutes and one download put a working model on your own hardware, still answering with the network switched off. The only real decisions are which of two free tools suits you and which model your memory can hold.

Updated 22 Sept 2026

The two tools worth your time are LM Studio and Ollama, and they differ in temperament more than in what they can do. LM Studio looks like a chat app: you browse models inside it, click download and start typing, and it shows you which ones fit before you commit to the download. Ollama is something you type — ollama run and a model name — and its real value arrives later, because it quietly serves an API on your own machine, so editors, scripts and coding agents can talk to your model instead of a hosted one. So: the app if you want to be chatting in fifteen minutes, the command if you expect to wire something onto it later.

Model names carry a number with a B after it, and it is worth decoding once. 9B means nine billion parameters, the internal numbers the model learned while it was trained. More of them usually means a better answer and always means a bigger file, and that file has to fit in your memory for the model to answer at a sensible speed. So the useful question is never which model is best; it is which of the good ones your machine holds.

The fit check answers that, and it is the only place on this site you need for it. Pick your machine from the list, or paste what your system reports about memory, and every model comes back with a verdict — it runs, it spills over into slower memory, or it will not fit — with the memory figure and an estimated speed beside it. Run that before you download anything, because a model that nearly fits is the worst outcome on offer: everything keeps working, just miserably.

One rule of thumb while you learn the shape of it: nine billion parameters sit inside 16 GB, and every step up in parameters asks for a step up in memory. Treat it as rough, because two things bend it. Compression is the first, since most people run a shrunk copy of a model, which is what makes any of this fit at all. The second is the mixture designs, where only a slice of the model works on each word: those need less memory than their name suggests. The picks below are three rungs of that ladder with their real verdicts one click away, and read the speed estimate beside the verdict as well — on an older laptop a model that fits comfortably can still answer at a pace you notice.

Expect the first answer to feel slower than a hosted service. That is your hardware being honest with you, and the fit check tells you roughly how slow before you spend the download. What you get in exchange: nothing you type leaves the machine, ever, and nobody is counting.

Where next: Check what your machine runs · Point your coding tools at it

  1. 01

    Check what your machine can run

    Open the fit check and pick your hardware, or paste what your system reports about memory. Every model comes back with a verdict for that machine and an estimated speed beside it. Write one model down before you open anything else; the three picks below are a reasonable shortlist to check first.

    checkpoint · You have one model name written down, with a runs verdict for your machine.

  2. 02

    Install one tool

    LM Studio if you want an app: browse, click download, start typing. Ollama if you want a command: one install, then ollama run and a model name in a terminal. Both run the same open-weight families, and their built-in catalogues differ a little, so pick your model inside whichever one you chose.

    checkpoint · The tool opens without errors.

  3. 03

    Download the model and say hello

    The first download is the slow part, because model files are gigabytes. Then ask it something real from your actual work, not a toy question, and watch the speed: words arriving at reading pace is what success feels like on ordinary hardware.

    checkpoint · A real question from your work got a usable answer, entirely offline.

  4. 04

    Prove to yourself it is private

    Turn the network off — flight mode, cable out — and ask again. This is the whole pitch made visible: a local model has no server to phone.

    checkpoint · It answered with the network off.

  5. 05

    Point something at it

    Ollama serves an API on your machine automatically; in LM Studio you switch the local server on from its developer view first. Either way, editors, scripts and coding agents can then talk to your model instead of a hosted one, which is the step that turns a demo into a tool.

    checkpoint · One app you already use is talking to your local model.

What to run it with

Questions people actually ask

What does the B in a model's name mean?+

Billions of parameters, which are the internal numbers the model learned during training. Treat it as a size label: more parameters usually means better answers and always means a bigger file. What matters to you is whether that file fits in your memory, and the fit check answers that for the machine you actually own.

Will it be as good as ChatGPT?+

No, and it is worth being clear about that before you spend an evening on it. The very largest models are not downloadable, and several that are ask for hardware most people do not have at home. What a mid-sized local model is genuinely good at is drafting, rewriting, summarising and answering questions about text you hand it, with nothing leaving the machine.

Do I need a graphics card?+

For the smaller models, no. An ordinary laptop with enough memory runs them, and Apple silicon Macs are unusually good at this because the whole machine's memory is available to the model. A card changes how fast the words arrive rather than whether they arrive, right up to the point where the model no longer fits inside it.

How much disk space does this take?+

Model files run from a couple of gigabytes to several tens of them, and each model page lists the size at every compression level. The download is the slow part of the whole exercise, and nothing after it is large. A model you dislike deletes in one click.

One install, one download, fifteen minutes, and a model that keeps answering after you pull the network out.