Writing captions by hand
On this page
Captions do not have to come from speech. There are two common reasons to type them.
There is no speech. A silent screen recording with captions explaining each step is a perfectly good video, and the recogniser has nothing to work with.
The video is in a language you are captioning differently, or you already wrote the script and it is more accurate than any recogniser will be.
Starting from nothing
On the Captions tab, with no transcript generated, you can add lines directly. Add one at the playhead makes one, you type into it, and you time it with the [ and ] keys while the video plays.
Pasting a script
There is a text box that takes the whole transcript as plain text, one caption per line. Paste the script you already have, and you get a caption per line, which you then time.
Clear empties it again.
What hand typed lines cannot do
Lines typed by hand carry no word level timing, because nothing measured when each word was said. They show as dimmed in the list.
The consequence is that editing the video by deleting words does not work on them: Cilevi does not know which part of the video each word belongs to, so it will not guess and cut your video in the wrong place.
Everything else works normally. They are styled, placed and exported exactly like recognised ones.