A good deal of the model work happening inside products has nobody watching it. A row of data arrives, a support ticket or a review or a signup, and the model reads it and hands back a category, a yes or no, a score out of five. It runs over three months of history once, then over new rows on a schedule, and the first most people hear of it is a tidier dashboard. Treat the model as a component, not as something you hold a conversation with, and nearly every decision that follows changes.
The first of those decisions is size, and the instinct to reach for the most capable model you can afford is worth resisting. A narrow task with a fixed prompt asks for something much more specific than general ability: it asks to be right about your six categories, on your writing, every time. The cards below start with a model small enough to live on a laptop and step up only as far as the work actually demands, and you can browse the same size class yourself among the small models carrying a permissive licence.
Which is why the benchmark that decides this is one you build. Label a few hundred rows by hand, hold a portion of them back, and measure each candidate against the rows it has never seen. That score describes your data, your categories and your prompt, which is more than any public ranking can offer, because none of them have seen your categories. Run it again on fresh rows every few months: the world your data describes keeps moving even when the model sits still.
There is a practical argument for downloading the model instead of calling someone else's. A classifier wired into a workflow should behave the same way next year as it does this afternoon, and a file on your own disk is the same file next year. A hosted model can be retired or quietly retrained while you are looking elsewhere, and the first sign of it is a drift in your numbers that costs a fortnight to chase. That is the most concrete reason open weights are worth something to a builder.
Then ask for a fixed shape, and check that you got it. Give the model the list of labels and tell it to answer with one of them, or with a small object of named fields, and validate every response before it reaches anything downstream. A classifier that occasionally writes a paragraph where a label should be is a bug three lines of validation will catch on the first run, and one that otherwise turns up as a corrupted column six weeks later.
As for where it runs, batch work is the most forgiving load there is, because nobody is waiting. A laptop grinding through rows overnight is a perfectly serious deployment, and the fit check will tell you whether the model you have chosen fits the laptop you own. A backlog too big for that fits the rung where you rent the hours the batch takes instead of a machine that sits idle between runs. And when the rows are customer data, running the thing yourself answers the privacy questions before they are put to you.
Where next: Small models with a permissive licence · What open weights mean · Estimating what a feature will cost