Can HeyGen Actually Create a Visual Explanation?
A timestamped review of whether one HeyGen quantum-superposition video teaches through pictures or mainly presents narration with related visuals.
The reviewed HeyGen video is a polished narrated presentation with some relevant graphics. It is not a complete visual explanation of quantum superposition.
That sentence is about one 49-second output made on August 1, 2026. It is not a claim that HeyGen is incapable of better work. The submitted prompt, creation workflow, settings, and interaction history were not preserved, so the sample cannot establish the product’s upper limit.
Watch the unedited HeyGen output.
The question this page answers
The comparison page asks which output performed better. This page asks a different question: if the sound were removed, would the visual sequence still teach the concept?
A visual explanation must show a relationship changing over time: before and after, cause and effect, part and whole, competing states, a quantity changing, or a process moving through steps. An avatar speaking, a subtitle, a relevant photograph, or a decorative animation may support a lesson without explaining it visually.
What HeyGen generated
The video alternates between a realistic presenter and short, dark-background cutaways. It uses large captions, repeated HeyGen watermarks, a spinning coin, a glowing particle, a small cat icon, a quantum-state grid, and a classical-versus-quantum panel. The transcript covers a coin analogy, measurement and collapse, Schrödinger’s cat, qubits, and quantum-computing speed.
The visuals are related to those subjects. Relevance is not the same as explanation.
Timestamped visual audit
| Time | Narration claim | What the picture shows | Does the picture teach it? |
|---|---|---|---|
| 0:00–0:04 | Something can be in two places at once; introduction to superposition | Presenter, a glowing dot, and an “Observation” badge | No. No two states or locations are represented. |
| 0:04–0:11 | A spinning coin is both heads and tails before it lands | A spinning coin labeled “Both States” | Partly. It supplies the analogy, but a classical spinning coin is not literally in both states. |
| 0:11–0:17 | A quantum system exists in all possible theoretical states | One glowing blue particle and captions | No. The states and their relationship are not visible. |
| 0:17–0:21 | Observation produces measurement or collapse | Presenter and captions | No. The video does not show a before-state changing into a measured outcome. |
| 0:21–0:32 | Schrödinger’s cat is alive and dead until the box is opened | Presenter with a small cat icon | No. There is no box, paired state, trigger, or reveal sequence. |
| 0:32–0:37 | A qubit can be zero, one, or both | A “Quantum Core” grid containing 0 and 1 nodes | Partly. The values appear, but superposition and measurement are not shown. |
| 0:37–0:44 | Quantum computers solve complex problems exponentially faster | A classical “AND” line beside purple branching lines | Weakly. The picture signals parallel branching but does not show an algorithm, a comparison, or the limits of the speed claim. |
| 0:44–0:49 | Superposition shows a mysterious and powerful universe | Presenter, “The Quantum World,” and a “Learn More” button | No. This is a closing card, not a visual conclusion. |
The exact visual failures
The central failure is at 0:17–0:21. Measurement is the point where the analogy needs a state change. The video shows the presenter instead. A viewer learns the term “collapse” from speech and captions but never sees what collapsed or how the result differs from the prior state.
The cat section repeats the same problem. A cat icon marks the subject, but an icon is not a model. The visual never establishes the sealed system, the two represented outcomes, or the observation event.
The last technical graphic arrives too late and is not connected to the earlier analogy. The 0/1 grid and branching lines look computational, but the video does not map coin states to qubit states or show why a quantum algorithm can exploit amplitudes and interference. The unqualified narration that quantum computers solve complex problems “exponentially faster” is also too broad. Speedups apply to particular algorithms and problem structures, not complex problems in general.
The video also spends a large share of its runtime on the presenter. That may improve social presence and brand fit, but it leaves less screen area and time for the mechanism. Repeated watermarks further compete with the learning material.
What works visually
The coin shot is the clearest explanatory moment. It gives the viewer a concrete object, motion, and two named outcomes. The classical-versus-quantum panel also attempts a direct comparison rather than showing unrelated footage.
The sequence is visually consistent. Captions are readable, cuts are clean, and the presenter is expressive and stable. Those strengths make the result easy to watch. They do not make its diagrams sufficient, but they explain why someone might prefer it as a presenter video.
Muted, full, and narration-only tests
The cleanest future evaluation uses three independent viewing conditions.
In the full pass, a reviewer sees sampled frames with the aligned transcript. In the muted pass, the reviewer sees the frames without transcript or audio. In the narration-only pass, the reviewer receives the transcript without frames. Each reviewer answers the same five concept questions and never sees the tool name.
The useful measures are:
- Comprehension: how many concept questions the full representation supports.
- Visual Lift: full-pass correct answers minus narration-only correct answers.
- Visual Carry: muted-pass correct answers divided by full-pass correct answers.
- Visual Explanation: whether non-text visuals encode causal, spatial, temporal, or quantitative relationships on a 0–10 scale.
For this output, the current 4/10 Visual Explanation score is a provisional manual judgment. No blinded reviewer sheet or learner study exists. The muted/full/narration protocol should be run before the score is treated as calibrated.
What HeyGen says the product can do
HeyGen describes Video Agent as a prompt-native system that writes the script, selects images, adds narration, creates transitions, and finalizes captions. Its getting-started guide also describes chat refinement and recommends moving a copy into AI Studio for precise manual changes.
The prompting guide says more context produces better results, recommends using another language model to refine the prompt, calls a full script the largest upgrade, and supports uploaded images, videos, PDFs, and documents.
This reveals the right interpretation of the sample. The unedited output tests what one unknown input and workflow produced. It does not test the best result obtainable with a full script, reference diagrams, multiple chat turns, asset replacement, and manual editing. Those extra steps also mean “one prompt to publish-ready explanation” should not be assumed from the marketing description alone.
What the research says to look for
There is no mature, highly cited independent study of HeyGen’s current Video Agent. The most defensible method is to apply established instructional-video research to the observed output.
Cynthia Brame’s review of effective educational video focuses on cognitive load, engagement, and active learning. The publisher reported 685 citations when checked on August 2, 2026.
Ibrahim and colleagues tested segmenting, signalling, and weeding. Their redesigned educational video improved transfer and structural-knowledge performance and reduced reported difficulty. The publisher reported 87 Crossref citations when checked.
Sung and Mayer found that instructive graphics improved recall while decorative and seductive graphics did not. This distinction is directly relevant here: a graphic can match the topic and still fail to show the relationship being taught.
Fiorella and Mayer’s instructional-video review reports that showing the instructor’s face by itself is not a reliable learning improvement. Wilson and colleagues’ four experiments found that learners could prefer instructor-present videos even when comprehension suffered. A polished avatar should therefore be scored as presenter quality, not counted as visual explanation.
Noetel and colleagues’ systematic review pooled 105 studies with 7,776 students and found positive average effects for video in higher education. A later critical reanalysis argues that those effects do not justify universal claims, especially for higher-order learning. The medium is not the mechanism: a video can help, but the pictures still have to do instructional work.
Answer
Can HeyGen create a visual explanation? This run shows that it can create visually relevant scenes and partial explanatory graphics. It does not show a complete visual explanation of quantum superposition. Most of the concept survives in the narration; too little survives in the pictures alone.
The next fair test is not another casual prompt. It is a logged, same-input, first-output experiment with an archived file, fixed concept questions, muted and narration-only passes, and independent reviewers.
Frequently asked questions
- Did the reviewed HeyGen video visually explain quantum superposition?
- Partly. It used a relevant coin analogy and a classical-versus-quantum graphic, but the central mechanism of measurement and collapse remained in the narration.
- Does a realistic AI presenter improve learning?
- Not necessarily. Research distinguishes viewer preference and social presence from comprehension. Presenter presence can help attention in some designs, but a face alone is not an explanatory visual.
- Is this a review of all HeyGen capabilities?
- No. It is a review of one unedited, vendor-hosted output whose prompt, workflow, settings, and interaction history were not preserved.