How to · Building with models

Using a small model as a classifier

Start smaller than feels comfortable. Sorting a support ticket into one of six buckets is a narrow, repetitive job, and the models that do it well are the ones you can keep on a laptop or rent for the hours a batch actually takes. Two hundred rows you have labelled by hand will settle the choice faster than any public ranking.

Updated 23 Sept 2026

A good deal of the model work happening inside products has nobody watching it. A row of data arrives, a support ticket or a review or a signup, and the model reads it and hands back a category, a yes or no, a score out of five. It runs over three months of history once, then over new rows on a schedule, and the first most people hear of it is a tidier dashboard. Treat the model as a component, not as something you hold a conversation with, and nearly every decision that follows changes.

The first of those decisions is size, and the instinct to reach for the most capable model you can afford is worth resisting. A narrow task with a fixed prompt asks for something much more specific than general ability: it asks to be right about your six categories, on your writing, every time. The cards below start with a model small enough to live on a laptop and step up only as far as the work actually demands, and you can browse the same size class yourself among the small models carrying a permissive licence.

Which is why the benchmark that decides this is one you build. Label a few hundred rows by hand, hold a portion of them back, and measure each candidate against the rows it has never seen. That score describes your data, your categories and your prompt, which is more than any public ranking can offer, because none of them have seen your categories. Run it again on fresh rows every few months: the world your data describes keeps moving even when the model sits still.

There is a practical argument for downloading the model instead of calling someone else's. A classifier wired into a workflow should behave the same way next year as it does this afternoon, and a file on your own disk is the same file next year. A hosted model can be retired or quietly retrained while you are looking elsewhere, and the first sign of it is a drift in your numbers that costs a fortnight to chase. That is the most concrete reason open weights are worth something to a builder.

Then ask for a fixed shape, and check that you got it. Give the model the list of labels and tell it to answer with one of them, or with a small object of named fields, and validate every response before it reaches anything downstream. A classifier that occasionally writes a paragraph where a label should be is a bug three lines of validation will catch on the first run, and one that otherwise turns up as a corrupted column six weeks later.

As for where it runs, batch work is the most forgiving load there is, because nobody is waiting. A laptop grinding through rows overnight is a perfectly serious deployment, and the fit check will tell you whether the model you have chosen fits the laptop you own. A backlog too big for that fits the rung where you rent the hours the batch takes instead of a machine that sits idle between runs. And when the rows are customer data, running the thing yourself answers the privacy questions before they are put to you.

Where next: Small models with a permissive licence · What open weights mean · Estimating what a feature will cost

What to run it with

Questions people actually ask

How many rows an hour should I expect?+

We do not hold a measured figure for your own laptop, and that is our gap, not something to estimate for you. Where we do hold throughput is on the hosted offers: some hosts report a rate measured on their own machines, and it sits on the model page beside those rows. Plenty of rows carry no figure at all. For your own machine the honest method takes ten minutes: time fifty rows, divide, multiply back up to the size of your backlog. Batch work is forgiving enough that the answer usually only has to beat "overnight".

Do I need a graphics card?+

For a model of a few billion parameters, an ordinary laptop manages, and nobody is waiting on a batch job anyway. A graphics card turns a night into an hour, which becomes worth it once you are rerunning the whole history every week. Run the model past the fit check, linked in the text above, before you assume either way.

Does the licence matter if this only runs internally?+

It can, and it is cheaper to read now than to unpick later. Some downloadable models carry conditions on commercial use or on passing the weights along, and "internal tool at a company that sells something" is exactly the case those conditions are written about. Every model page carries the licence line; the permissive ones say Apache or MIT and stop there.

What do I do when the categories change?+

Relabel a fresh sample and measure again. Changing the prompt is the cheap part of this job, and the expensive part is discovering six weeks later that a category you added in March has been quietly absorbed into the one next to it. New categories, new held-out set.

Should I fine-tune instead of prompting?+

Try the prompt first, because it is an afternoon and fine-tuning is a project. A good prompt with the label list and a handful of examples clears the bar for most sorting work. If your held-out score stalls somewhere short of what you need and the errors look systematic instead of random, that is the point where training on your own labels starts to earn its cost.

Two hundred rows you labelled yourself outrank every leaderboard here. Take the smallest model that clears the bar they set, then check it again in three months.