Subtitle Edit
About Subtitle Edit
Subtitles that are two seconds late are worse than no subtitles at all. Subtitle Edit exists to fix that, and about forty other things that go wrong with subtitle files, from a rip that arrived as pictures instead of text to a translation that runs off the bottom of the screen.
The window puts a video preview, a text grid and an audio waveform in front of you at once. You click a line, you see where it sits against the actual speech, and you drag it until it matches. That single arrangement is why this application has outlasted most of the paid tools in its category.
It is not the same job as Aegisub, which people reach for when subtitles need to be typeset, positioned, animated. This one is about timing and conversion, about correction, and about getting text out of places text does not want to leave.
The waveform is the whole point
Open a video and the audio appears as a waveform strip beneath it. Speech shows as visible blocks, silence as flat line, and dragging a subtitle’s start and end against those blocks is faster and more accurate than typing timecodes will ever be.
There is a spectrogram view as well, switched on in settings, and it earns its place on difficult material. Speech buried under music or traffic noise is often invisible in a waveform and perfectly legible as a band of frequencies in a spectrogram.
One thing catches new users out. The waveform is generated by an external encoder component, so if the encoding toolkit it leans on is not present, the strip stays empty.
Subtitle Edit offers to fetch it when a function needs it, and once it is there the waveform appears automatically on every video you open.
Three playback backends, and why you get a choice
Video preview in Subtitle Edit runs through one of several engines, selectable in settings. That sounds like a triviality until a file refuses to play or seeks inaccurately, at which point switching engines fixes in one click what would otherwise be an afternoon of codec archaeology.
There is also a re-encode helper for the worst offenders, which rewrites the video into something friendlier for subtitling work rather than asking you to fight the original.
Variable frame rate recordings from phones and screen capture tools are the usual candidates.
A sync tool for every kind of timing failure
This is the section that saves the most time, because subtitle timing goes wrong in several distinct ways and each has its own remedy.
A constant offset, where every line is late by the same amount, wants a fixed delay applied to the whole file. Progressive drift, where the gap grows through the film, is a frame rate mismatch and needs either a frame rate conversion or point synchronisation, where you match two known lines near the start and end and let everything between them be recalculated. Visual sync handles the same problem interactively, letting you set the first and last cue against the video.
The fourth option is the one people forget. If you have a correctly timed subtitle in another language, Subtitle Edit will align your file against it, transplanting the timings wholesale. When a translation exists but its timing is a mess, that is a ten-second fix rather than an evening.
Speech to text, and the proofreading it implies
Several speech recognition engines are supported, Whisper among them, and they run locally with nothing uploaded anywhere. You pick a model size, point it at the audio, and it produces a timed subtitle from scratch. Forced aligners can push that to word-level timestamps, and there is a text-to-speech path in the other direction for generating audio from subtitle text.
Be realistic about what arrives. Accuracy and speed both scale with your hardware and the model you chose, a large model on a modest processor being a job you start and walk away from.
What comes back needs reading, because names, technical terms and overlapping dialogue are where recognition falls down, and speaker separation needs extra setup and additional models on top.
Used properly it still changes the work fundamentally. Transcribing an hour of speech by hand takes most of a day. Correcting a machine transcript against the waveform takes a couple of hours, and Subtitle Edit is built for exactly that second workflow.
OCR, for subtitles that are pictures
Disc subtitles are images rather than text, which is why a rip so often leaves you with something no editor will open. Subtitle Edit is one of the few tools that reads them. This application reads Blu-ray SUP files, VobSub sub and idx pairs, XSub tracks inside AVI containers, and DVB or teletext subtitles out of transport streams, then runs OCR to turn the pictures into characters.
It also opens subtitle tracks sitting inside MKV and MP4 containers directly, so a muxed file does not need demuxing first. When it does, or when the subtitle stream needs pulling out of a disc structure, tsMuxeR handles that step cleanly.
Expect to check the output. OCR confuses the usual suspects, an uppercase I against a lowercase l, a one against an l, accented characters against their plain equivalents, and a pass through the correction tools afterwards is not optional. There is also a separate mode for working with the images themselves rather than converting them, useful when the source is stylised enough that recognition will never behave.
Fixing text in bulk
The fix-common-errors tool is the unsung feature. It runs a long list of checks over the whole file, catching missing spaces after punctuation, unnecessary periods, stray hyphens, capitalisation after ellipses, lines that exceed a sensible character count, and durations too short to read. Profiles let you keep different rule sets for different kinds of work.
Around it sit the rest of the text tools. Automatic line breaking, spell check against standard open dictionaries in many languages, plus find and replace with pattern matching, merging short lines and splitting long ones, and whole-file translation through online services with a review pass afterwards.
Batch conversion handles folders at a time, and there is a command-line converter for scripted work, which covers conversion, translation and OCR without the interface being involved.
For anyone processing a series rather than a film, that is the difference between a task and a routine.
Formats, and where the job hands off
Over three hundred subtitle formats can be read and written, which is the specification that makes this tool unavoidable in professional work. Alongside the familiar SubRip and WebVTT sit ASS and SSA, MicroDVD, the broadcast formats like EBU STL, cinema XML and TTML, and a long tail of things you will meet once and never again. Encoding is handled properly too, with UTF-8, Unicode and legacy single-byte encodings all supported.
What it does not do is put the subtitle back into a video container. Once the file is correct, muxing it into an MKV alongside the video and audio is a separate step, and MKVToolNix is the tool for it.
The other adjacent chore is naming. Players match subtitle files to videos by filename, so a folder of correctly timed subtitles with mismatched names helps nobody, and a renamer that pairs subtitle files to their videos closes that gap.
Conclusion
Subtitle Edit is the tool everyone working with subtitles ends up using, and the reason is that it covers the whole awkward middle of the job. Fansubbers correcting timing, accessibility teams producing captions, translators handed a file with unusable timings, anyone who ripped a disc and got pictures instead of text will find the specific fix they need somewhere in these menus.
What comes with that coverage is a program that hides nothing and explains little. The menu structure assumes you know what a frame rate mismatch looks like and which of four sync tools addresses it, and the automated features, meaning recognition and translation, produce drafts rather than results.
Learn the waveform, learn which sync tool matches which symptom, and treat everything automatic as a first pass. Do that and Subtitle Edit replaces a shelf of separate utilities with one window.
Pros & Cons
- Waveform and spectrogram views make timing a visual task rather than a numerical one
- Four distinct sync approaches, each matched to a different kind of timing failure
- Timings can be transplanted wholesale from a correctly timed subtitle in another language
- Local speech recognition produces a first draft without anything leaving the machine
- OCR covers Blu-ray SUP, VobSub, XSub and broadcast teletext sources
- Fix-common-errors applies dozens of checks across a whole file, with saveable profiles
- Over three hundred formats read and written, including broadcast and cinema specifications
- A command-line converter handles conversion, translation and OCR for scripted work
- The menu tree is enormous and nothing guides a newcomer through it
- The waveform stays empty until an external encoder component is present
- Recognition output always needs proofreading, particularly names and overlapping speech
- Speaker separation requires additional models and setup rather than working out of the box
- OCR results need a correction pass for confusable characters every time
- No muxing, so putting the finished subtitle into a container needs another program
Frequently asked questions
The waveform is drawn from audio extracted by an external encoder component. If it is missing, the strip stays blank, and the program will offer to fetch it the first time a function requires it.
That pattern means a frame rate mismatch rather than an offset. Use point synchronisation, matching one line near the start and one near the end, or convert the frame rate directly. A fixed delay will not help.
Yes, through several recognition engines that run on your own machine with nothing uploaded. Choose a model size to suit your hardware, then treat the result as a draft to correct against the waveform rather than a finished file.
That is one of its main uses. Image-based subtitles from discs and broadcast streams are imported and run through OCR, after which the text needs checking for the characters recognition typically confuses.
It connects to online translation services and can process an entire file at once. Quality varies by language pair, and idioms and context are where machine translation reliably falls apart, so the review pass matters.
Yes. A separate command-line converter ships alongside it, covering format conversion, translation and OCR, which suits processing a whole series in one scripted pass.