Generating a transcript
On this page
With nothing selected, open the Captions tab and press Generate transcript.
What is doing the work
Whisper, a speech model, running on your own processor. The audio does not leave the machine, there is no account and no per minute charge, and it works with no internet connection once the model file is on disk.
It knows 99 languages out of one model file, so there is nothing to download per language.
The three sizes
Standard, 264 MB. Fast, and good for English.
Better quality, 574 MB. Slower, and much better for Persian, Arabic and most languages that are not English.
Best quality, 1.1 GB. Slowest, and the most accurate on casual or fast speech.
The first time you choose one it is downloaded, checked against a known fingerprint so a broken download cannot be used, and kept under %LOCALAPPDATA%\Cilevi\models. Deleting that folder costs nothing but the download.
Pick Standard for English. Pick one of the other two for anything else, and expect the larger ones to take longer: this is running on your processor, so the trade is real.
Language
Auto works out which language it heard and tells you what it decided. It can get this wrong on short or noisy speech, and a misidentified language produces confident nonsense rather than an obvious error.
So if you know the language, pick it. The choice is remembered for next time.
The prompt
Edit prompt takes a few lines of text that help with words the model has never heard: your product name, a person's name, a piece of jargon. One per line. This is the fix for a name that comes out spelled wrong every time.
Also available
There is a second engine: the speech recogniser Windows itself ships. It needs no download at all. It is less accurate and knows fewer languages, and it is there for the case where you cannot spare the disk space or the wait.