New: Boardroom MCP Engine!

Ready to put this into action?

Get the complete AI Integration PlaybookPractical AI implementation guide — prompt engineering, workflow automation, and ROI frameworks.

Article 060 · Part 6

Build a Video from Script to Captions

Make the narration, visuals, captions, and description tell the same supported story.

By Randy Salars · Published

On this page
  1. Define a modest video brief
  2. Write the evidence-backed narration first
  3. Label the role of every shot
  4. Assemble a small asset list
  5. Create captions from the final audio
  6. Review meaning across all channels
  7. Check disclosure requirements for the destination
  8. Watch the actual final export
  9. Adapt without losing the qualification
  10. For students: keep the production decisions visible
  11. Practice: assemble a 60–90-second package

Make the narration, visuals, captions, and description tell the same supported story.

The narrator says that nine of twelve respondents selected a favorable survey answer. The video shows a cheering crowd and a headline that reads, “Everyone loved it.” The words are careful; the finished message is not.

Video combines several channels of meaning. Images, editing, music, text, and captions can reinforce the script or quietly contradict it. A reliable production process checks each channel and then watches the whole export.

Start with one clear promise. For a short educational video, decide what viewers should understand by the end and what they should be able to check for themselves.

Define a modest video brief

This original production example teaches viewers to identify the group behind a percentage. The planned duration is 75 seconds, within a 60–90-second target. The main format is a wide educational video, with a possible separate short version later.

The source is synthetic: 24 confirmed registrants, 18 attendees, 12 attendee survey respondents, and nine respondents selecting “welcoming.” The other three response choices are unspecified, and six attendees did not respond. No real event, recording, or survey is being represented.

The intended result is a viewer who asks, “Seventy-five percent of whom?” The video does not need a dramatic setting, a celebrity-style presenter, or a claim about improved event outcomes.

This article provides a script and storyboard plan. No footage, narration, captions, or video export has been produced or timed for it.

Write the evidence-backed narration first

Use the source packet to check every statement before choosing visuals. Keep each sentence understandable when heard once. Put important qualifications into the spoken explanation instead of relying on a tiny footnote.

The storyboard below includes the complete working narration. The planned shot durations total 75 seconds; they must be adjusted after recording. Numbers and timing are production assumptions until the actual performance is available.

Planned timeWorking narrationVisual and on-screen textEvidence and media status
0–10 seconds“Two labels can show seventy-five percent and describe different groups. Here is a fictional example.”Two editable labels, each reading “75%,” followed by “Of whom?”Original explanatory graphics; label example as fictional
10–25 seconds“Twenty-four people reserved places at a reading session. Eighteen attended. Eighteen out of twenty-four is seventy-five percent of the registrants.”“18 attendees / 24 registrants = 75%” with clear group labelsChecked arithmetic from synthetic packet
25–43 seconds“Twelve attendees answered a survey. Nine selected ‘welcoming.’ Nine out of twelve is also seventy-five percent, but this group is the survey respondents.”“9 welcoming responses / 12 respondents = 75%”Checked arithmetic; no claim about all attendees
43–58 seconds“Six attendees did not respond. We do not know how they would have answered. The other three respondents’ choices are not specified either.”Separate labels for “6 did not respond” and “3 other responses: unspecified”Preserve distinct missing-information categories
58–70 seconds“Before using a percentage, ask how many people are counted and which group they belong to. Put that group in the label.”Two questions: “How many?” and “Out of which group?”Instruction derived from the worked example
70–75 secondsNo narration; allow a short reading pause.“Check the group before trusting the percentage.” Source note: “Synthetic teaching example.”Original closing card; no real-event footage

The visual does not need to depict exactly twenty-four people. Editable numbers and simple shapes are easier to verify. If icons encode a count, ensure the count is correct and explain what each icon represents.

Label the role of every shot

A shot can be documentary evidence, an illustration, a simulation, or an abstract explanation. Record which role it has.

In this example, every visual is an original explanatory graphic. A generated crowd would add little and could imply that a real event had occurred. If you use a conceptual scene for atmosphere, label it and keep it separate from the numerical evidence.

For a historical video, an actual source photograph and a generated reconstruction serve different purposes. For a scientific video, a diagram and a recorded experiment are different kinds of evidence. The viewer should not have to guess which they are seeing.

Use this prompt:

Turn this verified script into a shot-by-shot storyboard. Identify each shot as documentary, illustrative, simulated, or explanatory graphics. Preserve counts, groups, and uncertainty. Describe what each visual adds and flag any image that could imply a stronger claim than the narration. Do not invent source footage or completed production work.

Assemble a small asset list

For the percentage example, the required assets are the reviewed narration, editable number graphics, a source card, and any licensed music chosen for a specific purpose. No music is necessary to teach the lesson.

Keep the source table and calculations with the project. Record rights and provenance for any external asset. Do not assume that a short clip or a background track is available for unrestricted use because it is easy to download.

Choose project settings appropriate to the destination and your source media. Check frame size, aspect ratio, frame rate, audio format, and caption support using the editor and platform’s current guidance. Avoid unnecessary conversions that introduce scaling, motion, or synchronization problems.

If you plan a vertical version, design it separately enough that labels remain readable. Cropping the wide frame can remove the denominator—the very detail the lesson is teaching.

Create captions from the final audio

Captions should match what is actually spoken, not an earlier script version. Correct numbers, names, negations, and speaker information. Include meaningful non-speech audio when it contributes to understanding.

W3C explains captions as synchronized text for speech and relevant non-speech audio and stresses that automatic captions need correction. A transcript can help with access and production, but untimed text alone is not the same as captions on the video. W3C WAI: Captions and Subtitles

For the sample, inspect these meaning-changing errors:

Incorrect captionRequired correction
“Ninety selected welcoming”“Nine selected ‘welcoming’”
“75 percent of the attendees”“75 percent of the registrants” or “respondents,” matching the specific spoken sentence
“We know how they would have answered”“We do not know how they would have answered”

Time the cues against the actual recording. Give viewers enough time to read them and avoid covering essential labels. Do not claim a caption file is synchronized simply because its timestamps add up to the planned duration.

Review meaning across all channels

Watch once with sound and once without it. The silent pass helps reveal whether the visible labels preserve the message. Then listen without watching to see whether the spoken explanation remains understandable.

Compare the title, thumbnail, description, narration, graphics, and captions. A title such as “Proof Everyone Loved Our Event” would contradict the entire example, even if every frame inside the video were carefully labeled.

A suitable working title is “Before Trusting a Percentage, Check the Group.” A matching description is: “A fictional reading-session example shows why the same percentage can refer to different groups. All counts are synthetic teaching data.”

The thumbnail could ask “75% OF WHOM?” using editable text. It should not use a fabricated crowd as evidence of attendance or a real person’s face as an implied endorsement.

Check disclosure requirements for the destination

Platform requirements can change, so consult the current official guidance before uploading. YouTube’s guidance, for example, requires disclosure for realistic content that is generated or meaningfully altered in specified ways, while distinguishing some non-realistic or minor edits. Its examples include synthetic scenes, altered depictions of real people or places, and AI-generated music. Evaluate the actual content against that guidance rather than assuming that all AI assistance is treated identically. YouTube: Disclosing Use of GenAI Content

A platform label does not replace the source note that tells viewers a dataset is fictional. Nor does a disclosure grant permission to use another person’s likeness or an unlicensed asset.

The current example remains a manuscript production plan. No upload or publishing action is implied.

Watch the actual final export

Open the file you will deliver and watch it from beginning to end. Inspect clipped text, unintended blank frames, abrupt cuts, audio distortion, caption timing, and the final source card. Check the exported duration and the beginning and ending of both picture and sound.

Do not rely only on the editing timeline. Rendering can reveal problems that were not obvious in the preview. If you upload a permitted test version, inspect the platform’s processed playback and caption behavior too.

Keep a short review record with the file version, checks performed, defects found, and corrections. An unresolved caption error should not disappear from the record because the overall video feels polished.

Adapt without losing the qualification

A short clip might use just one percentage comparison, but it still needs the group label and fictional-data context. Do not remove the denominator to make room for a larger number.

For a different audience, change the vocabulary or pace while preserving the claim.

For students: keep the production decisions visible

Follow the assignment’s rules for generated media, voice use, and collaboration. Keep the storyboard and contribution record so you can explain the production decisions. Identify what you wrote, created, selected, edited, and checked, along with any assistance you received.

A video portfolio should identify a plan as a plan and an exported, reviewed video as an exported, reviewed video. The distinction is part of professional communication.

Practice: assemble a 60–90-second package

Create a brief, source packet, working script, storyboard, asset list, caption plan, title, and description. If you produce the video, time the captions to the actual audio and review the complete export. Record what changed after watching it.

Completion check: Narration, visuals, captions, title, and description communicate the same supported meaning. Illustrations and synthetic data are identified. Required permissions and platform disclosures are addressed for the actual media used. A completed-video claim is supported by a real export and a documented review.

Get the AI Dispatch

Weekly insights on ai & technology — delivered to your inbox. No spam, unsubscribe any time.

Want to choose specific topics? Customize your interests