A recorded lecture arrives as one continuous file. Somewhere in it is a slate the lecturer read to the camera, four sentences addressed to whoever was operating it, a false start they immediately corrected, eleven pauses longer than anyone wants to sit through, and a noise floor made of an air conditioner and a preamp.
None of that is teaching. All of it is in the file. And every single thing we do afterwards — every caption, every animated cue, every B-roll cutaway, every frame of the presenter keyed onto a background — is timed against that file and carries its audio.
So the question is not whether to clean it. The question is whether you clean it before you build on it, or after — and after means building it twice.
Why these two run first
This is not a preference about tidiness. It is the shape of the dependency graph.
Every effect in the pipeline is driven by a word-level timeline — a file that says which word was spoken at which millisecond. The captions read it. The animated cues are bound to it by phrase lookup. The B-roll cutaway windows start and end on its phrase boundaries. The presenter's mouth is matched to it within two frames. And every one of them carries the same voice track.
Cut a single second out of the middle of the lecture after that timeline exists and everything downstream of the cut moves. Every caption, every cue, every cutaway window, every sync point. There is no partial re-render; the timeline changed, so the render changed.
The order between the two is causal too, and it surprises people. Studio Sound profiles the room noise from the recording itself rather than from a preset — and it has to profile the take that actually ships. Run it first and you build a noise profile out of material that is about to be cut, describing a room the finished video never contains.
Pass one: Video Cleanup
It audits the file, transcribes it to the word, removes what was said for the crew rather than the learner, corrects a picture the camera got wrong, and measures everything later passes will need.
Audit. Variable frame rate, rotation flags, baked-in letterbox bars, interlacing, a dead audio channel, dropped frames. Every one of these is invisible until something downstream breaks in a way that looks like a completely different bug.
Transcribe to the word. A 25-minute take has too many small artifacts to catch by ear. A timestamped transcript is also the only way to cut precisely on a word boundary instead of halfway through one.
Remove what isn't teaching. Slates and take markers. Production talk — slide callouts, spoken stage directions, asides to someone off-camera. Retakes, in full rather than trimmed. Filler words, lightly. Doors, coughs and phones that land in a gap.
Trim the dead air, keep the breath. Anything under 1.2 seconds is a real pause and stays. Anything longer comes down to about half a second — not to zero, because narration with every pause removed sounds rushed and synthetic.
Correct the picture. Black level, white balance, exposure drift — measured across the whole take, not from one frame. Correct, not graded: style is the effect's job later.
Survey the frame. The real greenscreen colour. How far the presenter moves, sampled across the entire take. How much headroom the framing has. Measuring once here means no later pass has to re-derive it — and means a crop chosen later cannot clip a gesture it never saw.
On a recent 27-minute undergraduate lecture that came to 58.7 seconds removed across 40 spans — 27:14 down to 26:16. Eight false starts, one doubled word, thirty-one over-long pauses. Zero slates and zero production talk, which is itself worth telling a customer: that recording process is already clean.
Pass two: Studio Sound
It strips the background noise a room leaves behind, un-muddies the voice, gives back the top end a clip-on microphone loses, and levels the whole thing to a steady broadcast loudness that does not drift across twenty-five minutes.
Everything it does is decided from a measurement, because nobody can hear their own processing after twenty minutes with it and a lecture recording has no reference to compare against. On the same lecture:
| Measurement | As recorded | After |
|---|---|---|
| Signal-to-noise ratio | 19.4 dB | 29.5 dB |
| Noise floor | −35.9 dBFS | −45.5 dBFS |
| Integrated loudness | −18.2 LUFS | −16.4 LUFS |
| True peak | +0.5 dBFS — over full scale | −1.4 dBTP |
| 250–800 Hz (boxiness) | +14.3 dB | −1.3 dB |
| 6–16 kHz (air) | −19 dB | +7 dB |
The true peak line is the one worth dwelling on. A recording can sit below full scale on every sample and still overshoot it between samples — which is inaudible in the raw file and distorts the moment it is encoded for delivery. It is the kind of defect that only ever shows up in the version the student watches.
Half of both passes is deciding not to do something
This is the part that does not look like work, and it is most of the value.
Not broken
- De-clipping. 53 samples sat at full scale. Grouped into runs they were 26 isolated bursts of one to five samples — 0.1 ms. Inaudible. Repairing them replaces real signal with interpolation for no gain
- Hum notching. No mains harmonic more than 2 dB above its local floor. A clean electrical environment; a notch filter would only have removed voice
- De-essing. Sibilance measured normal. A de-esser on an already-dull source makes it duller
Left in deliberately
- "the newspaper's real consequence. Customers" — the transcriber's error at a chunk boundary. She said "the newspaper's real customers"
- "the app quietly log locks" — same cause. She said "logs"
- Two more apparent retakes that were the same artifact
- One genuinely ambiguous phrase, flagged rather than cut, because no script existed to check it against
Those four would have been cut by a pass that trusted its first transcript. They are not filler — they are sentences the lecturer said, and removing them would have quietly damaged a good take in a way nobody would notice until a student hit a hole in the argument. Every candidate retake now gets re-transcribed in a narrow window and confirmed before it is removed.
A cleanup pass that removes everything it flags is not careful. It is just fast.
What neither pass can do
Worth saying plainly, because the alternative is discovering it after a purchase order.
Studio Sound cannot restore frequencies that were never captured. If the microphone did not hear above 6 kHz, the air shelf is lifting whatever little is there — it helps, and it is not the same as a microphone that heard them. A lavalier clipped outside clothing rather than under it would beat our entire processing chain, and costs nothing. Neither pass can remove a cough that overlaps speech, separate two people talking at once, or undo sustained clipping. And neither can supply a point the lecture never made.
We would rather hand that list over at the start than send back something polished and hope nobody looks closely. It is the same reason we ask people to send us their worst recording rather than their best one.
Why they stay at the top of the pipeline
Because the cost of doing them late is not additive, it is multiplicative. A second removed after the captions are burned in is a re-render of the captions, the animated cues, the B-roll windows and the presenter sync. A noise profile built before the cuts describes a room that no longer exists in the file. A crop chosen before anyone measured how far the presenter moves clips her hand the first time she gestures, somewhere nobody happened to check.
Every one of those is cheap to get right at the front and expensive to fix at the back. That is the whole argument. It is not craftsmanship for its own sake — it is the order that means we only build the video once.
Send one recording and we will run both passes on it and tell you exactly what we found — including the categories that came back empty. Send a lecture →