Transcribe our videos and podcasts once, then subtitle them in every language we publish in
We publish a long-form video and a podcast episode every week, and each one goes out in five languages. The recording is the easy part. What holds a release up is the captions.
What the job looks like today
- Someone runs the audio through a transcription tool and cleans up the transcript by hand.
- A translator takes the cleaned transcript and returns five documents.
- Someone else pastes each translation into a subtitle editor and re-times it, line by line, because the timings do not survive the round trip.
Step three is the one that costs us a day per episode. The words are already correct by then — we are paying for the clock.
What we need
- One upload that produces the transcript and every language from it, rather than a transcript we then have to carry somewhere else.
- Speakers separated. Half our output is two or more people in conversation and a transcript that does not say who is talking is not usable as a transcript, let alone as a subtitle.
- Timings that carry across languages. The translation of a line belongs in the same slot the original occupied. This is the whole request.
- SRT and VTT out, because that is what our video host and our podcast player each want.
What matters more than accuracy
A word we have to fix is a small cost. A line that sits on screen for one second or runs to three rows is one a viewer turns the captions off over — and once they are off for one episode they stay off.
So: split lines for reading speed, keep them short enough to read at a glance, and let us see the result before it goes anywhere. We would rather review five languages in an editor for twenty minutes than trust five files we never opened.
Volume
- 4 videos + 4 podcast episodes a month, 20 to 90 minutes each
- 5 languages: English, Turkish, German, Spanish, French
- A back catalogue of roughly 200 episodes with no captions at all, which we would work through if the per-episode cost were close to zero