The phrase "second brain" covers two different ambitions. The modest one is a searchable pile of everything you said and heard: meetings, voice memos, the idea you dictated at a bus stop. The ambitious one links those notes together so that next March, when a client asks what was agreed, the answer surfaces itself. You can reach the modest version in an afternoon. The ambitious one is a habit you keep up for months, and you build it yourself: no model on this site will do it for you.
What models genuinely solve is the first three stages: turning speech into text, turning long text into short text, and answering a question across a folder of notes. All three run on your own machine, which matters more here than usual, because a meeting recording is mostly other people's words and transcribing it at home means the audio never leaves the room it was recorded in.
Start with the transcription model, because everything downstream inherits its mistakes. A meeting is the hard case for this kind of model: several voices, some of them talking across each other, and a microphone in the middle of a table instead of in front of a mouth. Models that read a clean studio recording almost perfectly lose several times as many words in that room, so the accuracy figure you see quoted in a headline is measured on the wrong thing. The board to judge on is the recorded-meetings board, which scores every speech model we track on real recordings made in a room, lowest error at the top.
Read it before you pick, because the ordering surprises people. The famous name sits a long way down it. The rows at the very top are closed models, marked so in the board's own column: there is nothing to download, and for the one in front we list no host to rent from either. The two best-placed models you can download are Granite Speech from IBM and Cohere Transcribe, both about the size an ordinary laptop swallows without complaint, both free to download, and both a long way above Whisper on this particular audio. The cards below name each model and link to its page, where the download size and the fit verdict for your machine stay live.
What the famous name still buys is the software around it. Whisper has free apps on every platform, it covers far more languages than either model above it, and you can have a transcript from it in twenty minutes without opening a terminal. The two named above are recent releases, so for now they are scripts and command lines. If a first transcript tonight matters more than the last word of accuracy, start with Whisper and move when the errors begin to annoy you. If this is going to be a nightly habit, start with the best-placed downloadable model on that board.
Speed is the other axis, and it stops mattering quickly. The speed board ranks the same models on how much faster than real time they run, and every model named here clears an hour of audio in far less than an hour. That gap only decides anything if you are transcribing live or clearing a backlog of years. The audio tab of the catalogue gathers every transcription model we list with its speed, languages and licence in one place. Read its accuracy column as a separate question, though: it averages all the English recordings the leaderboard tests, of which a meeting room is only one, so models that separate in a room sit almost level there. For the room itself, stay on the meetings board.
The last stage is the one everybody skips, and it is the one that makes the other four worth doing, so here is how it actually goes. You do not point a model at the folder. You search the folder, the way you search anything, for the word you half remember. Then you paste the two or three notes that came back into the model you already run and ask your question against those. That is retrieval done by hand, it needs nothing you have not already installed, and it beats the automatic versions for a long time because you can see precisely what the model was given before it answered.
Arithmetic is what makes it work this way. A month of meetings runs well past any context window you have, and handing a model more than it can hold makes it find less. Every tool that claims to answer questions across your documents is doing this same search on your behalf and hiding the result, which is fine when it picks the right note and confusing when it picks the wrong week.
When the folder outgrows hand-searching, and it will somewhere in the low hundreds of notes, the thing you want is a local tool that keeps an index of the folder for you. AnythingLLM is built for exactly this and runs entirely on your machine; Open WebUI, if you already use it as a front end for Ollama, will take a document set too. Insist on one thing from whichever you choose: it must show you which note an answer came from. These tools find notes by matching the shape of your words against the shape of theirs, and that is confidently wrong often enough that an answer about your own meetings is only worth having when you can open the note behind it.
We have no guide of our own for that indexing step yet. The gap is ours, and it is on our list.
Here is the whole pipeline, in five stages, each closing on a checkpoint you can verify before moving on.