Use B-roll for talking-head video when the viewer needs to see a place, object, process or piece of evidence that speech alone cannot show clearly. Use a designed graphic when the viewer needs to compare, remember or understand structured information. Keep the speaker on screen when expression, trust or personal delivery is the main information.
The decision should follow meaning. It should not follow a rule such as "change the picture every few seconds."
The short answer
Ask one question at each sentence:
What does the viewer need to see in order to understand or trust this claim?
The answer usually points to one of four visual treatments:
| Viewer needs | Best starting visual |
|---|---|
| Human expression or personal authority | Speaker on screen |
| A real object, place or action | B-roll or screen recording |
| A number, list or comparison | Designed graphic |
| A source or exact interface state | Screenshot or source view |
If the current speaker frame already carries the information, do not add another visual merely to create activity.
A-roll, B-roll and designed graphics
Adobe defines B-roll as supplemental footage that supports the main clips. It can establish a scene, smooth a transition or add meaning. A-roll is the primary footage that carries the main action, dialogue or interview.
In a talking-head video:
- A-roll is usually the person speaking to camera.
- B-roll is supporting footage such as a product close-up, location, process, archive clip or screen recording.
- Designed graphics are constructed visual explanations such as metric cards, comparisons, lists, charts and diagrams.
- Captions represent spoken language and should not be treated as a replacement for every other visual.
B-roll and graphics can both cover a jump cut, but hiding a cut is not enough reason to use either one. The inserted visual should also support the sentence.
Keep the speaker when the person is the evidence
Stay on the speaker when viewers need to evaluate confidence, emotion, sincerity or personal experience.
Examples include:
- A founder explaining why a decision was made.
- A customer describing a problem in their own words.
- An instructor introducing a difficult distinction.
- A creator delivering a joke or personal reaction.
- A speaker qualifying a claim or acknowledging uncertainty.
Cutting away during these moments can weaken the connection between the words and the person saying them. A tighter crop may provide emphasis without replacing the face.
Use B-roll when reality needs to be seen
B-roll is strongest when it answers "what does this look like?" or "how does this happen?"
Show the object being discussed
If the speaker names a product control, material or physical detail, show the real object. The image should match the exact version being described.
Show a process
Use a screen recording, over-the-shoulder view or close-up of the action. Keep the sequence long enough for the viewer to follow the important step.
Establish place and context
An exterior, workspace or event view can orient the viewer before detailed explanation begins. Do not reuse an attractive location shot if it suggests the wrong place or time.
Cover a necessary dialogue cut
Relevant B-roll can bridge two sections of spoken audio and avoid a distracting jump. Adobe notes that supplemental footage can help smooth transitions and remove unwanted frames without discarding a whole shot.
The visual still needs a relationship to the words. Generic typing hands, city traffic or coffee pouring rarely explains a software workflow.
Use designed graphics when information has structure
Speech is linear. Some ideas are easier to understand when their relationships can be seen at once.
Numbers and metrics
Show the exact number, unit and context. A large number without a denominator, period or source can mislead even when it matches the spoken word.
Lists and sequences
Use a concise list when the viewer needs to remember several items. Reveal timing should follow the spoken sequence instead of showing every line before it is mentioned.
Comparisons
A two-column comparison can clarify a decision when both sides use the same criteria. Avoid creating a comparison that favors one option through unequal wording or missing limitations.
Processes and systems
A flow or structure diagram helps when the speaker describes dependencies, stages or ownership. Keep the number of nodes small enough to read at video size.
Quotations and sources
Show a short, accurate excerpt or source title when evidence matters. Do not crop away context that changes the meaning, and do not display text that the viewer cannot read in the available time.
Choose the visual with a decision table
| Spoken moment | Keep speaker | Use B-roll | Use graphic |
|---|---|---|---|
| Personal opinion | Usually | Only for relevant context | Rarely |
| Product demonstration | Brief introduction | Yes | For labels or steps |
| Exact number | Before or after | Only if it shows the source | Usually |
| Before and after | For framing the result | Yes | For criteria or summary |
| Abstract workflow | For explanation and trust | Sometimes | Usually |
| Emotional story | Usually | Only if authentic and relevant | Rarely |
| Source-backed claim | For delivery | Show the source | Summarize carefully |
This is a starting point, not a template. One moment can combine the speaker and a small graphic if both remain readable.
Time the visual to the spoken claim
A supporting visual should enter close enough to the relevant sentence that the connection is obvious.
Use this review sequence:
- Find the exact words that introduce the subject.
- Place the visual when the viewer first needs it.
- Keep it visible for the action or reading task.
- Return to the speaker when personal delivery matters again.
- Listen and watch from the previous sentence through the next one.
Do not cut away before the visual has a purpose, and do not leave it on screen after the subject changes.
Generated B-roll needs stricter review
Generated imagery can be useful for concepts that were not filmed, but it introduces additional risks:
- A product may have the wrong shape, interface or branding.
- A place may look real while representing no real location.
- A historical or scientific detail may be incorrect.
- A generated person may imply an endorsement or event that never happened.
- Visual continuity may change between shots.
Label generated visuals when context requires it. Never use generated footage as documentary evidence. For factual demonstration, prefer real footage, screenshots or diagrams built from verified information.
Rights and source checks
Before using any supporting asset, record:
- Where it came from.
- What license or permission applies.
- Whether people, logos or private information appear.
- Whether the crop changes the original meaning.
- Whether attribution is required.
Stock availability is not the same as permission for every use. Generated media also does not remove responsibility for privacy, likeness and intellectual-property review.
Common visual mistakes
Adding movement without information
Frequent cutaways can make the edit feel busy while reducing comprehension. Keep a stable frame when the sentence itself deserves attention.
Using B-roll that only matches one keyword
A clip of a generic robot does not explain AI video editing. Choose the actual interface, operation or result being discussed.
Turning every sentence into a card
Graphics lose hierarchy when everything receives the same treatment. Reserve designed elements for information the viewer needs to compare, remember or verify.
Covering the speaker at the wrong moment
Do not cut away from a reaction, qualification or emotionally important line simply because B-roll is available.
Forgetting caption space
Check supporting visuals with captions enabled. A useful product detail can become invisible when subtitles occupy the same screen zone.
A practical workflow
First complete the spoken structure using the talking-head first-cut workflow. Then mark only the sentences where the viewer needs additional visual information.
For each mark:
- Write the visual job in plain language.
- Choose real footage, screen capture or a designed explanation.
- Verify factual and rights-sensitive details.
- Place it at the corresponding transcript moment.
- Review the actual frame with captions.
- Remove the visual if it does not improve understanding.
An AI talking-head video editor can prepare framing, captions and designed blocks in one composition. The editor still decides whether each visual earns its time on screen.
Frequently asked questions
How much B-roll should a talking-head video use?
There is no fixed ratio. Use enough to show what speech cannot, establish context and cover necessary cuts. Keep the speaker visible when their expression and authority matter.
Is B-roll only video footage?
The term usually refers to supplemental footage, but photos, screen recordings and archive material can serve a similar supporting role. Designed graphics are better treated as a separate category because they construct an explanation.
Can stock footage replace custom B-roll?
Only when it accurately represents the subject and you have the right to use it. Custom product, interface and process footage is usually more specific and trustworthy.
Should B-roll cover every jump cut?
No. Clean jump cuts are normal in creator video. Cover a cut when the jump distracts or when a relevant visual adds information.
When should I use a graphic instead of B-roll?
Use a graphic for structured information such as numbers, comparisons, lists and systems. Use B-roll for real objects, places and actions.
Start with one section of a real recording. Mark the exact sentences that need proof, demonstration or structure, then add only those visuals in Pireel Studio.


