Most people meet transcription through a website that wants an upload, and conclude that getting words out of audio is something you rent. It is not. This is the one job where a model small enough to sit on an ordinary laptop does work you would otherwise pay for, and it does it without the recording ever leaving the room.
That last part is why it is worth the afternoon. A meeting is mostly other people's words. Uploading it makes a decision on their behalf that they were not asked about, and doing the same work locally removes the question entirely.
The steps below are the short version of a first run. The stance is the part worth arguing with: start with the app rather than the model, and take the middle-sized one. Nearly everyone who tells you local transcription is too slow started at the top of the ladder on a machine that could not carry it.