Studio
Studio turns a recorded capture into video. Not a screen recording with a voiceover on top, but a produced clip with characters, scripted narration, art direction, and your real product screens cut into it.
It exists because the two things teams want from product video pull against each other. A screen recording is accurate but nobody watches it. A polished marketing video gets watched but goes out of date the moment the UI moves, and re-shooting costs more than the feature did. Studio’s answer is to make the recording the source of truth and generate everything around it, so a re-record regenerates the film.
⚠️ The engine never draws a product UI, so “no capture” means “no screen” — never an invented one.
This is worth stating plainly, because it is not how generative video normally behaves. Ask a general video model for “the billing settings page” and it will produce something confident, plausible and wrong: right shape, right colours, invented labels. For a mood film that is a stylistic choice. For documentation it is a support ticket, and in a regulated industry it is considerably worse.
Studio will not do it. Every frame showing your product is a frame someone actually recorded. Where a step was never captured, the engine works around the gap — it cuts to a character, holds a title card, or narrates over something else — but it will not draw the screen.
That last row is the one to internalise. A reel is not the video of a capture — it is a video of it. One SmartDoc session holds many reels.
A single recording of your onboarding flow can produce a 9:16 clip for social, a 16:9 walkthrough embedded in the docs, a sixty-second teaser, and a chaptered tutorial — different characters, pacing and screen density in each — without anyone recording anything twice.
This also changes video maintenance. When the UI moves, you re-record once. Every reel built from that source can be regenerated. The expensive, manual, always-skipped part of keeping video current collapses into one recording session.
Studio is where the MotionVideo lane lives, and it accepts work from the other two.
MotionVideo is the lane to use when you have images but no recording — a design mockup, an exported frame, a diagram, a screenshot someone sent you.
Studio is strongest where content is procedural: a flow with steps, a setup with an order, a feature with a beginning and an end. That is what the capture pipeline sees and what the guide structure encodes.
💡 The test: if you would demonstrate it by sharing your screen, Studio will do well. If you would explain it on a whiteboard, write a page with a diagram instead.
Video generation is the most expensive thing on the platform by a wide margin, and Studio is built around that rather than hiding it.
The pipeline deliberately pauses before the expensive stage. Character animation is the slow, costly part, so the engine plans, scripts, draws and voices first, then holds at awaiting_review and waits for you.
Usage rolls into AI Usage & Cost alongside everything else, so video spend sits next to the rest of your AI usage rather than in a separate bill nobody opens.
💡 If you are evaluating Studio, generate one reel end to end and look at what it cost before planning a library. That number is usually the deciding input, and it is better discovered on reel one than reel thirty.