How it works

Most lectures don't need forty animations.

Animated B-roll cuts away from the lecture to an illustration for a few seconds, then cuts back. The hard part is not drawing it. The hard part is knowing when not to.

There is an obvious way to add animation to a recorded lecture, and it is wrong. You decide the video needs, say, one cutaway a minute; you find the nearest plausible sentence to each mark; you draw something for it. Twenty minutes in, twenty animations out.

What you get is a video where roughly half the graphics are illustrating nothing in particular, because the schedule asked for a picture at 14:00 and the lecturer happened to be saying "and this brings us to the next point." Students read that instantly, even if they could not tell you why. The animation stopped meaning anything, so they stopped looking at it.

We do it the other way round. Read the lecture, find the moments that actually contain a picture, and let those decide the count. Sometimes that is thirty. Sometimes it is nine.

The unit is the phrase, not the second

Every lecture we produce is transcribed to the word before anything is drawn, and those words are grouped into caption phrases — the same phrases that drive the on-screen captions. A cutaway is always one of those phrases, or two to four adjacent ones. It begins and ends exactly where a phrase begins and ends.

That single rule removes the worst failure this effect has. A cutaway that starts a beat early clips the front of a word; one that runs a beat long leaves a picture hanging over the start of the next sentence. Both read as a mistake. Tying the window to the phrase grid makes them impossible rather than unlikely.

Pan & Zoom Parallax Build-On Floating Glow 0:00 0:30 1:00 1:30 2:00
Two minutes of a real lecture. Each grey tick is one caption phrase — 47 of them, averaging 2.5 seconds. The four pink windows are the cutaways, and every edge lands on a tick boundary. 11 of the 47 phrases sit inside a cutaway: 23% of the segment.

What earns a cutaway

Once the transcript is on the table, the question for each phrase is narrow: does this name something a picture can hold? In practice four kinds of phrase survive being drawn, and everything else is better left on the lecture's own view.

Cut away

A picture adds something

  • A process. "Somebody built it, somebody pays for it, somebody gets something back"
  • A list. "Advertising, subscription, sponsorship and data"
  • A comparison. Horizontal versus vertical integration — one owner, many outlets against one company, the whole chain
  • A concrete concept. A five-rupee newspaper that costs more than five rupees to print
Stay on the lecture

A picture is just decoration

  • Connective tissue — "so that", "in other words", "and here is the trap"
  • A sentence already carried by the deck's own typography
  • Anything abstract enough that the drawing would be a stock metaphor
  • A long narrative stretch. A story is already vivid; cutting away competes with it
  • The final phrase and the silence after it. Never end on a graphic

Four ways to move, and they are not interchangeable

A still illustration can be animated four different ways, and picking one house treatment and repeating it is how a set of cutaways starts to feel mechanical. The treatment should follow what the sentence is doing.

PAN & ZOOM

For a single payoff. The camera eases across an oversized still and pushes toward the thing the sentence is building to — a number, a checkmark, a price tag. Use it when the composition has one obvious place to end up.

PARALLAX

For scene-setting. The camera holds still; background, glow and foreground drift at different speeds. Depth without anything visibly panning. Best when there is no single focal point, or as relief after a stretch that already has a lot of its own motion.

BUILD-ON

For an enumeration. Nothing moves; elements arrive one at a time, each cued to the exact word that introduces it, with a connecting path drawing alongside. Items that have not arrived yet are drawn as dashed placeholders, so the frame never looks broken while it fills.

FLOATING GLOW

For a definition that has to sit. A static crop; the only motion is ambient — elements bobbing on out-of-phase sine waves, the glow breathing. Use it when a long caption needs to be read and camera movement would compete with reading.

So how many, for a twenty-minute lecture?

This is the question we get asked, and the honest answer is that we cannot tell you before reading the script — but we can show you the arithmetic on a real one.

Take a twenty-six minute undergraduate lecture on the business of media. Transcribed, it comes to 569 caption phrases averaging 2.49 seconds each. A cutaway of two to three phrases therefore runs five to seven and a half seconds.

Now read it for illustratable moments. There are about thirty: the cost stack behind a cheap newspaper, the loop where content buys attention and attention buys content, the four revenue models, the incentive split between an ad-funded outlet and a subscription one, a brand paying more to reach a known nineteen-year-old than a stranger, horizontal versus vertical ownership, an iceberg the lecturer describes out loud, one cricket match monetised two ways, a five-point recap. Thirty genuine moments in twenty-six minutes — roughly one every fifty seconds.

569
caption phrases in the lecture
~30
moments that actually contain a picture
12–18%
of runtime spent in cutaway

Thirty cutaways at five to seven seconds is about four minutes — twelve to fourteen percent of the runtime, or fifteen to eighteen if the richest handful are allowed to run longer. That is the number the content supports. To fill a third of the video you would need forty to sixty cutaways, one every twenty-five seconds, and the last twenty-five of them would be padding with a drawing on top.

The percentage is a ceiling and a sanity check. The concept inventory is the constraint that actually binds.

Density is also uneven on purpose. A two-minute story gets one or two cutaways. A stretch that enumerates and compares can carry one every twenty or thirty seconds. Between any two, the lecture's own view returns for at least one full phrase, so the video keeps coming back to itself rather than becoming a slideshow with narration.

The parts nobody notices unless they are wrong

A cutaway replaces the entire frame, captions included. So the captions have to be redrawn on top of the illustration, from the same word timings — which means they land in exactly the same place, at exactly the same size, with the same word highlighted, on both sides of every cut. Get that wrong by a pixel and every cut flickers.

Behind the redrawn captions sits a soft legibility plate, because an illustration has no guaranteed contrast against caption ink the way a controlled background does. Its size is computed once per phrase from the widest weight any word will take, not per frame — otherwise the plate grows and shrinks as words change weight, which is worse than having no plate at all.

And every illustration is drawn larger than the frame, so that no element is caught crossing an edge part-way through a camera move. We check that by sampling the border of every frame across the whole cutaway, not by looking at the composition once and assuming.

Two things we got wrong on that lecture

Both were caught by those checks rather than by eye, which is rather the point of having them.

The legibility plate was originally a flat black at 50% opacity — the standard answer, and the wrong one here. On a light background with dark caption ink, a black plate moves the backdrop toward the text rather than away from it. Measured, it came out at 2.97:1 contrast, under the 3:1 floor for large text. The plate now shifts with the scene — near-white on a light ground, black on a dark one — and measures 13:1.

And the first Pan & Zoom had three small screen glyphs and their connecting lines caught crossing the left edge three seconds into the push-in. The fix was not to loosen the zoom and lose the payoff; it was to fade the periphery out before the crop ever reaches it.

Neither would have been visible in a still. Both would have been visible in the video, to everyone, forever.

Send one recording and we will show you where we would cut away in it, and where we would not. Send a lecture →

See it on your own lecture.

One recording, a produced cut back within 24 hours, free — including an honest note on which moments were worth animating and which were not.