I Built a Claude Skill That Makes Videos (HyperFrames + HeyGen)
I gave Claude Code one sentence and got back a finished explainer video. Here is the skill I built, what it renders well, and the part it still gets wrong.
I typed one sentence into Claude Code and walked away. When I came back there was an MP4 on my desk that I had not touched: researched, scripted, narrated, animated, rendered.
No timeline. No After Effects. No motion designer invoice.
I have been using this on my own channel for months. The B-roll in my last 20 videos came out of it. This is the build, start to finish, including the parts that did not work.
The Piece That Makes This Possible#
The tool underneath all of this is HyperFrames. It is an open-source framework from HeyGen that turns Claude Code into a rendering engine.
Here is the part that matters, and the part most people miss. HyperFrames is not a generative video model. It does not invent footage and hope it looks right. Claude writes real HTML and CSS, and HyperFrames renders that to video.
That distinction is the whole ballgame:
- Text is legible, because it is actual text, not a model's guess at letterforms
- Brand colors are exact hex values, not approximations
- The same prompt renders the same video twice, so you can iterate instead of re-rolling
- It runs locally, so there is no per-clip credit burn
What The Skill Actually Does#
A Claude skill is a folder of markdown that teaches the agent a job. Mine handles the whole chain so I do not have to think about any of it.
Give it a topic and it researches it, writes a beat-mapped script, generates the voiceover, builds every scene in one locked visual language, burns in word-synced captions, and renders the MP4.
The phrase I keep coming back to: the skill is not a shortcut for people who already know the workflow. The skill is the workflow.
That matters because the failure mode of AI video is not bad rendering. It is inconsistency. Generate scenes one at a time and you get eight clips that each look fine and nothing that looks like one video. The skill enforces a single visual language across every scene, which is the thing you cannot prompt your way to by hand.
Where HeyGen Comes In#
Motion graphics with a voiceover is most of a video. It is not all of one. At some point a person needs to be on screen.
That is the one thing the agent cannot produce on its own, and it is where HeyGen slots in. Through the HeyGen MCP, Claude calls out for the talking-head A-roll and drops your AI avatar straight into the render.
I made a clone of myself and put it in front of the graphics. In the video I hold the real me and the avatar side by side and ask which is which. I genuinely could not tell.
A few practical notes from doing it:
- Your clone needs about 15 seconds of footage minimum. Two minutes or more gets noticeably better
- Name the exact avatar and outfit you want in the prompt. The agent will match it
- You do not need the MCP to use HyperFrames. It only matters if you want avatars
The Honest Verdict#
HeyGen sponsored this video. You know the drill, I do not hold back.
HyperFrames is genuinely good at the structural work and the price is impossible to argue with. Where it falls down is fine detail. Out of the box you get output that is 80 percent of the way there and then stalls on the last 20, which is exactly the 20 that separates something you would publish from something you would not.
That gap is why I built the skills rather than just prompting it directly.
The weakest part of any first pass is the voiceover. Every time. Listen to it critically and ask for a rewrite before you accept the render. Two other things that cost me time: install HyperFrames first and add the MCP second, and say explicitly whether you want horizontal or vertical, because otherwise you get both.
The Resend launch video it produced came out so clean I did not want to put an avatar on it at all.
Both Skills, Free And Paid#
I am giving the explainer skill away. It is the one that makes Vox-style shorts, and it ships with a reference project that renders in about 19 seconds so you can confirm your setup works before building anything.
The second one is the production version. Same pipeline, with the judgment added: brand-kit sourcing, ad sets in 16:9, 9:16 and 1:1 from one run, hook variants for paid testing, and an 11-gate quality system built out of every note a client ever sent back.
hyperframes doctor and it will tell you what is missing.Should You Bother#
If you make explainer content, yes. The math is not close. A motion designer is hundreds of dollars per video and days of turnaround. This is your existing Claude subscription and about twenty minutes.
If you need live action, real people, real locations, this is not that and it is not pretending to be.
Everything I used is in the video, and both skills are linked above. Tell me what you want me to build with it next.
Related Reading#
- How to make AI videos, the step-by-step version of this workflow with the real costs broken down
- AI Clone System: scale to 10 videos a week without filming, more on building and using an avatar
- AI 3D animation: 4 professional styles in under 10 minutes, for when you want generated visuals instead of rendered ones
- AI Story Videos: full pipeline from research to final edit, the manual version of this pipeline
- 30 videos in 30 days: what the data actually showed, the volume experiment this workflow came out of
Moe shares tool walkthroughs and lessons from real projects. Mechanical engineer, then venture capital, now building AI tools for creators and small businesses. More about Moe
Get new videos in your inbox
Weekly AI workflows. No fluff.
No spam. Unsubscribe anytime.