Psst ships with five engines and marks the one that fits your Mac. This is the reasoning behind that mark, so you can override it with your eyes open.
The short version
| Your Mac | Use | Why |
|---|---|---|
| 8 GB, any chip | Apple Speech | Nothing to download, nothing to hold in memory, good English. |
| 16 GB | Whisper small.en | About 500 MB on disk, a few hundred MB in RAM, fast, quiet. |
| 24 GB and up | Whisper large-v3-turbo | About 1.5 GB on disk, the best accuracy you can run locally today. |
| M Pro, M Max, M Ultra | Whisper large-v3-turbo | The Neural Engine and GPU have room; run it all day. |
Apple Speech
This is the engine built into macOS. It runs on the device on Apple silicon, needs no download, and is surprisingly good at clean English. It is the right default for a Mac with 8 GB, and a fine choice on any Mac when you want zero setup. Its limits: fewer languages than Whisper, and it drifts on technical vocabulary.
Whisper small.en
An English-only Whisper, about 500 MB. It punches above its size for dictation because it only has to know one language. On a 16 GB Mac it leaves plenty of room for the browser and the editor you are dictating into.
Whisper large-v3-turbo
The big one, about 1.5 GB, multilingual, and the most accurate model Psst can run without a server. "Turbo" means the decoder was trimmed so it runs several times faster than the original large-v3 with almost no loss. On a Mac with 24 GB or more, and especially on an M Pro or M Max, this is the one to leave on all day.
What "memory" means here
Apple silicon shares memory between the CPU, GPU and Neural Engine. A model loaded for transcription takes its share from the same pool as your apps. That is why the recommendation is by total memory, not by chip generation. A base M1 with 16 GB will run small.en happily; an M4 with 8 GB should stay on Apple Speech.
Switching later
Nothing is locked. Settings shows every engine with its download size. Download two if you like, and pick per recording. The transcripts are plain text either way.