How it works

The two passes nobody asks for.

Universities ask us for animation, for captions, for branding. In two years nobody has asked for cleanup. It is the first thing we run, and the thing that decides whether any of the rest works.

A recorded lecture arrives as one continuous file. Somewhere in it is a slate the lecturer read to the camera, four sentences addressed to whoever was operating it, a false start they immediately corrected, eleven pauses longer than anyone wants to sit through, and a noise floor made of an air conditioner and a preamp.

None of that is teaching. All of it is in the file. And every single thing we do afterwards — every caption, every animated cue, every B-roll cutaway, every frame of the presenter keyed onto a background — is timed against that file and carries its audio.

So the question is not whether to clean it. The question is whether you clean it before you build on it, or after — and after means building it twice.

Why these two run first

This is not a preference about tidiness. It is the shape of the dependency graph.

Every effect in the pipeline is driven by a word-level timeline — a file that says which word was spoken at which millisecond. The captions read it. The animated cues are bound to it by phrase lookup. The B-roll cutaway windows start and end on its phrase boundaries. The presenter's mouth is matched to it within two frames. And every one of them carries the same voice track.

Cut a single second out of the middle of the lecture after that timeline exists and everything downstream of the cut moves. Every caption, every cue, every cutaway window, every sync point. There is no partial re-render; the timeline changed, so the render changed.

01 Video Cleanup what is in the take audit the file cut slates, retakes, dead air correct the picture survey the frame 02 Studio Sound what it sounds like measure everything subtract the room shape the voice level to broadcast loudness words_clean.json the timeline every effect times against studio_audio.wav the voice track every effect carries Full Animated Lecture Presenter + Animated Presenter + Captions Animated / Image / Video / Stock B-Roll Objectives, Takeaways & Credits every one of these reads the two files on the left
The two passes produce two files. Everything after them is a consumer of those two files, which is why they are not optional and not last.

The order between the two is causal too, and it surprises people. Studio Sound profiles the room noise from the recording itself rather than from a preset — and it has to profile the take that actually ships. Run it first and you build a noise profile out of material that is about to be cut, describing a room the finished video never contains.

Pass one: Video Cleanup

It audits the file, transcribes it to the word, removes what was said for the crew rather than the learner, corrects a picture the camera got wrong, and measures everything later passes will need.

00

Audit. Variable frame rate, rotation flags, baked-in letterbox bars, interlacing, a dead audio channel, dropped frames. Every one of these is invisible until something downstream breaks in a way that looks like a completely different bug.

01

Transcribe to the word. A 25-minute take has too many small artifacts to catch by ear. A timestamped transcript is also the only way to cut precisely on a word boundary instead of halfway through one.

02

Remove what isn't teaching. Slates and take markers. Production talk — slide callouts, spoken stage directions, asides to someone off-camera. Retakes, in full rather than trimmed. Filler words, lightly. Doors, coughs and phones that land in a gap.

03

Trim the dead air, keep the breath. Anything under 1.2 seconds is a real pause and stays. Anything longer comes down to about half a second — not to zero, because narration with every pause removed sounds rushed and synthetic.

04

Correct the picture. Black level, white balance, exposure drift — measured across the whole take, not from one frame. Correct, not graded: style is the effect's job later.

05

Survey the frame. The real greenscreen colour. How far the presenter moves, sampled across the entire take. How much headroom the framing has. Measuring once here means no later pass has to re-derive it — and means a crop chosen later cannot clip a gesture it never saw.

On a recent 27-minute undergraduate lecture that came to 58.7 seconds removed across 40 spans — 27:14 down to 26:16. Eight false starts, one doubled word, thirty-one over-long pauses. Zero slates and zero production talk, which is itself worth telling a customer: that recording process is already clean.

Pass two: Studio Sound

It strips the background noise a room leaves behind, un-muddies the voice, gives back the top end a clip-on microphone loses, and levels the whole thing to a steady broadcast loudness that does not drift across twenty-five minutes.

Everything it does is decided from a measurement, because nobody can hear their own processing after twenty minutes with it and a lecture recording has no reference to compare against. On the same lecture:

MeasurementAs recordedAfter
Signal-to-noise ratio19.4 dB29.5 dB
Noise floor−35.9 dBFS−45.5 dBFS
Integrated loudness−18.2 LUFS−16.4 LUFS
True peak+0.5 dBFS — over full scale−1.4 dBTP
250–800 Hz (boxiness)+14.3 dB−1.3 dB
6–16 kHz (air)−19 dB+7 dB

The true peak line is the one worth dwelling on. A recording can sit below full scale on every sample and still overshoot it between samples — which is inaudible in the raw file and distorts the moment it is encoded for delivery. It is the kind of defect that only ever shows up in the version the student watches.

Half of both passes is deciding not to do something

This is the part that does not look like work, and it is most of the value.

Measured, then skipped

Not broken

  • De-clipping. 53 samples sat at full scale. Grouped into runs they were 26 isolated bursts of one to five samples — 0.1 ms. Inaudible. Repairing them replaces real signal with interpolation for no gain
  • Hum notching. No mains harmonic more than 2 dB above its local floor. A clean electrical environment; a notch filter would only have removed voice
  • De-essing. Sibilance measured normal. A de-esser on an already-dull source makes it duller
Looked like defects

Left in deliberately

  • "the newspaper's real consequence. Customers" — the transcriber's error at a chunk boundary. She said "the newspaper's real customers"
  • "the app quietly log locks" — same cause. She said "logs"
  • Two more apparent retakes that were the same artifact
  • One genuinely ambiguous phrase, flagged rather than cut, because no script existed to check it against

Those four would have been cut by a pass that trusted its first transcript. They are not filler — they are sentences the lecturer said, and removing them would have quietly damaged a good take in a way nobody would notice until a student hit a hole in the argument. Every candidate retake now gets re-transcribed in a narrow window and confirmed before it is removed.

A cleanup pass that removes everything it flags is not careful. It is just fast.

What neither pass can do

Worth saying plainly, because the alternative is discovering it after a purchase order.

Studio Sound cannot restore frequencies that were never captured. If the microphone did not hear above 6 kHz, the air shelf is lifting whatever little is there — it helps, and it is not the same as a microphone that heard them. A lavalier clipped outside clothing rather than under it would beat our entire processing chain, and costs nothing. Neither pass can remove a cough that overlaps speech, separate two people talking at once, or undo sustained clipping. And neither can supply a point the lecture never made.

We would rather hand that list over at the start than send back something polished and hope nobody looks closely. It is the same reason we ask people to send us their worst recording rather than their best one.

Why they stay at the top of the pipeline

Because the cost of doing them late is not additive, it is multiplicative. A second removed after the captions are burned in is a re-render of the captions, the animated cues, the B-roll windows and the presenter sync. A noise profile built before the cuts describes a room that no longer exists in the file. A crop chosen before anyone measured how far the presenter moves clips her hand the first time she gestures, somewhere nobody happened to check.

Every one of those is cheap to get right at the front and expensive to fix at the back. That is the whole argument. It is not craftsmanship for its own sake — it is the order that means we only build the video once.

Send one recording and we will run both passes on it and tell you exactly what we found — including the categories that came back empty. Send a lecture →

Find out what is actually in your file.

One recording, a produced cut back within 24 hours, free — with the full cleanup and audio report, honest about what the source would not give up.