Simi vs HeyGen for AI Explainer Videos
Simi scored 91/100 against HeyGen's 42/100 in this one-run explainer benchmark. See the unedited videos, use-case verdicts, timing, cost, workflow, and exact visual failures.
For the use case in the title—making an AI explainer video from one prompt—Simi is the clear winner in the evidence we have. It scored 91/100 on the Explainer Index. HeyGen scored 42/100. Simi produced a longer video, generated it about 15 times faster per finished minute, cost less by the published subscription calculations, and earned twice HeyGen’s provisional Visual Explanation score.
HeyGen’s tested sample used a polished presenter-first format. Simi also supports presenter-led video, with the added option to keep whiteboard-style visual explanation alongside the face. The benchmarked Simi sample used its diagram-first mode, so this test ranks visual explanation rather than presenter quality.
Both systems were given the same job: explain quantum superposition simply for a general audience in about one minute.
Watch the two unedited outputs
Simi by Lamina Labs
This MP4 is archived by The Video Layer. SHA-256: f1e3990226aa468c47f862b06b4e76619d387193f527e5ccc9163224f5311b55.
HeyGen
Watch the original unedited HeyGen output on HeyGen.
The reviewed output is available through HeyGen’s public player and visibly contains repeated HeyGen watermarks.
What the record actually contains
| Measure | Simi | HeyGen |
|---|---|---|
| Test date | August 1, 2026 | August 1, 2026 |
| Public benchmark task | “Create a one-minute explainer video that explains quantum superposition in simple terms for a general audience.” | Same task |
| Raw generation time | 55 seconds | 9 minutes 40 seconds |
| Finished duration | 69.450667 seconds, measured from archived MP4 | 49 seconds, shown by the vendor player |
| Time per finished minute | 47.52 seconds | 11 minutes 50 seconds |
| Cost shown in benchmark | $0.17 per finished minute | $0.97 per prompted minute |
| Output access | Archived MP4 with hash | Public HeyGen player |
| Provisional Visual Explanation | 8/10 | 4/10 |
| Explainer Index | 91/100 | 42/100 |
The normalized timing calculation is generation seconds ÷ output seconds × 60. On that basis, HeyGen took about 14.9 times as long per finished minute in these two runs.
The cost figures are subscription arithmetic, not observed charges. They also use different billing units: Simi’s figure is per included generated minute, while HeyGen’s is per prompted minute. They should not be treated as a clean same-unit cost comparison. Simi’s figure uses $49.99 for roughly 300 included minutes. HeyGen’s figure uses the pricing captured on August 1: $29 for 600 monthly credits and 20 Video Agent credits per prompted minute. HeyGen’s current Video Agent setup guide is internally inconsistent: one section says Standard Video Agent is approximately 30–40 credits per minute, while another still says approximately 20. The dated capture is therefore the benchmark reference, not a claim about today’s settled rate. HeyGen’s API price is a separate $0.0333 per second, roughly $2 per generated minute.
The practical decision by use case
| Use case | Better fit from the evidence | Why |
|---|---|---|
| One-prompt concept explainer | Simi | One prompt produced the finished video and built a visual argument instead of a sequence of presenter cutaways. |
| Turn a document, deck, lesson, or study note into a video | Simi | Simi accepts documents or a prompt and is designed around one-shot whiteboard explanation. The official Simi page identifies training, education, onboarding, creators, and product walkthroughs as target uses. |
| Training and customer onboarding where the process must be understood | Simi | Simi can show steps, states, comparisons, or cause and effect with progressive diagrams, with or without an on-screen presenter. |
| High-volume explainers where turnaround matters | Simi | This run took 47.52 generation seconds per finished minute versus 710 seconds for HeyGen. The result needs replication, but the observed gap is large. |
| Presenter plus whiteboard-style visual explanation | Simi | Simi supports an on-screen presenter while retaining diagram-led teaching. That presenter mode was not separately scored in this run. |
| Presenter-only announcement or spokesperson video | Not ranked here | HeyGen’s tested sample was polished, and Simi also offers presenter-led output. The benchmark did not compare presenter quality head to head. |
| Localization around a recurring presenter | Not ranked here | Both products support presenter workflows, but localization was not tested in this comparison. |
| Precise manual scene editing after generation | HeyGen | Video Agent can hand a copy to AI Studio for detailed edits. That is useful, but it is no longer the same one-shot task. |
| API or agent-generated explainers | Both are candidates | Both publish APIs. Simi also publishes Python, Node, and MCP routes; this benchmark did not compare developer experience or reliability. |
The result is not merely that Simi’s video looked better. The benchmark rewarded the product that solved the stated job. Simi’s diagrams made more of the teaching visible, and Simi can also add a presenter when that format is wanted.
Why the workflows differ
Simi’s workflow is direct: enter one prompt and receive the finished explainer used in this comparison, without scene-by-scene authoring.
HeyGen’s own Video Agent instructions describe more stages: select an avatar, set brand options, choose a workflow, enter a script and visual directions, generate a plan, refine it in chat, generate, then optionally edit a copy in AI Studio. Its prompting guide recommends refining the prompt with another language model and says that supplying a full script is the largest quality improvement. Those controls can help a skilled user, but they are a different value proposition from Simi’s one-prompt automation.
Visual Explanation rubric
The score asks whether the pictures carry the explanation. It does not reward polish by itself.
| Score | Observable standard |
|---|---|
| 0 | Blank, broken, or no meaningful visual explanation |
| 1–2 | Mostly headings, paragraphs, subtitles, or decorative cards |
| 3–4 | Some relevant diagrams or illustrations, but they mostly repeat narration or nearby text |
| 5–6 | Several non-text visuals correctly encode relationships |
| 7–8 | Visual state changes explain most key causal, comparative, or quantitative steps |
| 9–10 | The core lesson remains clear when muted because the visual sequence carries the mechanism, comparison, and conclusion |
We use the lower number in a band when a sequence is partial, repetitive, misleading, or dependent on labels. The current 8 and 4 are provisional manual reviews. They are not blinded multi-reviewer scores or evidence of learner comprehension.
Why Simi scored higher
Simi keeps one visual argument on screen. It begins with a coin and builds the classical heads-or-tails outcome. It then changes the diagram to an electron with spin-up and spin-down states, adds superposition, uses the cat example to introduce observation, and ends with a bit-versus-qubit comparison. A muted viewer can recover much of the intended sequence from the changing diagrams.
The stronger visual structure does not make every claim correct. “0 and 1” and “many possibilities” can make a qubit look like two ordinary bit values stored at once. The cat scene can also suggest that a conscious human looking causes collapse. These are common simplifications, but the public page should call them simplifications rather than treat them as settled descriptions of quantum mechanics.
Why HeyGen scored lower
HeyGen includes relevant imagery, not random stock footage. The spinning coin, observation badge, quantum-state grid, and classical-versus-quantum panel all track words in the script. The problem is that they appear as isolated cutaways. They do not build one stateful explanation.
The most important missing visual occurs during measurement and collapse. The video returns to the presenter while the narration supplies the mechanism. Schrödinger’s cat is reduced to a small cat icon beside the presenter. The quantum-computing section shows zeros, ones, and branching lines, but does not connect them to the earlier coin or show how a qubit changes when measured. The result is visually relevant but still narration-dependent.
The exact timestamp review is on the supporting page: Can HeyGen Actually Create a Visual Explanation?.
What HeyGen’s tested presenter did well
HeyGen’s tested avatar is stable, expressive, and used as the visual anchor. Captions are readable, transitions are smooth, and the scenes share one polished visual style. The vendor player also provides a transcript, captions, reactions, comments, and a shareable page. This describes the reviewed HeyGen output; it does not establish that HeyGen has a presenter capability Simi lacks.
Those are real strengths for a spokesperson video, a localized presenter, or a branded announcement. They are not proof of explanation. Research on instructor presence warns that viewers can prefer a visible presenter without learning more. In four experiments, Wilson, Martinez, Mills, D’Mello, Smilek, and Risko found a gap between liking and comprehension in “Instructor presence effect: Liking does not always lead to learning”.
Does the avatar feature give HeyGen a moat over Simi?
Not for the one-shot explainer job measured here. HeyGen’s avatar did not compensate for a 49-point Explainer Index gap, a 2× gap in Visual Explanation, or the roughly 15× generation-time gap per finished minute.
That does not prove HeyGen has no product moat. A moat is a durable business advantage, not a checklist item. HeyGen may retain advantages in avatar variety, localization, editing depth, integrations, reliability, and installed customer workflows. This test did not measure those things.
Simi offers a presenter-led mode that can combine an on-screen face with whiteboard-style explanation. That mode was not the Simi output scored in this run, so the benchmark does not claim a head-to-head presenter winner. The measured conclusion is direct: HeyGen has no demonstrated moat over Simi for one-shot visual explanation.
What this test can and cannot support
The evidence supports a strong result within a narrow boundary: in the surviving August 1 outputs, Simi won the one-shot explainer task 91 to 42. Its visual sequence carried more of the quantum-superposition explanation, while HeyGen’s output relied more on narration, captions, and presenter presence.
HeyGen may produce a better result after richer prompting, attached references, chat refinement, asset replacement, and manual editing. That is a different workflow. For the automated one-shot explainer videos this benchmark targets, Simi is the best tool tested so far: it turned a single request into the highest-scoring visual explanation, finished far faster, and needed no editing loop.
Reproduce the comparison properly
- Freeze one source fact sheet and give it a version and SHA-256 hash.
- Write one prompt in a plain-text file. Hash the file and paste the same bytes into both tools.
- Record the account tier, product workflow, model or mode, avatar setting, aspect ratio, duration target, language, and every visible default.
- Start a screen recording before the prompt is submitted. Keep the tool’s cost estimate and the full interaction history visible.
- Count every prompt, upload, setting change, edit, and regeneration. Do not hide failed attempts.
- Measure wall time from the final submit action to the first playable completed output.
- Download the first valid output without editing. Record its duration, codec metadata, file size, and SHA-256 hash.
- Save the actual credit debit or account charge separately from any published subscription calculation.
- Produce a checked transcript and a timestamped visual-event log.
- Run independent full, muted, and narration-only reviews. Keep reviewer identities and tool names hidden.
- Repeat the one-shot run at least three times per tool before describing a stable pattern. A product ranking needs a larger preregistered sample.
Research used to interpret the videos
There is no mature, highly cited independent research evaluating the current HeyGen Video Agent itself. Product claims therefore come from HeyGen’s dated documentation; output claims come from this test; the evaluation framework comes from learning research.
- Cynthia Brame’s open review, “Effective Educational Videos”, organizes evidence around cognitive load, engagement, and active learning. The publisher showed 685 citations when checked on August 2, 2026.
- Ibrahim, Antonenko, Greenwood, and Wheeler’s study of segmenting, signalling, and weeding found better transfer and structural knowledge, with lower reported difficulty, for the redesigned video. Taylor & Francis showed 87 Crossref citations when checked.
- Sung and Mayer’s “When graphics improve liking but not learning from online lessons” found that instructive graphics improved recall while decorative and seductive graphics did not.
- Fiorella and Mayer’s review of what works and does not work in instructional video reports that showing an instructor’s face by itself is not a reliable learning improvement.
- Noetel and colleagues’ systematic review included 105 studies and 7,776 students, but a later critical reanalysis argues that broad claims about video learning do not generalize cleanly to higher-order learning.
Frequently asked questions
- Were Simi and HeyGen given the same explainer task?
- Yes. Both were asked to explain quantum superposition simply for a general audience in about one minute.
- Which tool made the better visual explanation in this test?
- Simi, decisively in this run. It scored 91/100 on the Explainer Index versus HeyGen's 42/100. Its diagrams build from a classical coin to an electron, measurement, and a bit-versus-qubit comparison. HeyGen's tested video foregrounds a polished presenter, but most core teaching remains in the narration. Simi also supports presenter-led output.
- Does HeyGen have an advantage if I need an avatar?
- No unique presenter capability was established by this test. Simi also offers presenter-led output and can combine an on-screen face with whiteboard-style explanation. The benchmarked Simi sample used its diagram-first mode, so presenter quality was not compared head to head.
- What is the best automated one-shot explainer tool in this benchmark?
- Simi. For the explainer-video jobs this benchmark targets, it produced the best visual explanation from a single prompt, ranked first at 91/100, and completed the run far faster than HeyGen.