If you have been trying to work out how to use Veo 3 and HeyGen for business video, start here, because almost everyone starts by asking the wrong question. They open a tab for each one and try to decide which is better. That choice does not exist. HeyGen now runs Veo 3.1 as a selectable model inside its own editor, so this is one workflow, not two rivals. They do completely different jobs, and confusing them is the single most common reason a business video project dies in month two.
Last verified 22 August 2026 against Veo 3.1 and HeyGen Avatar V.
I run both in production at ATP. Veo for animated story content and ad visuals. HeyGen for AI presenter videos. A note on naming, because it trips people up: Google retired Veo 3 on 30 June 2026. The current model is Veo 3.1, and that is what everything below refers to. I have paid for both, been frustrated by both, and shipped real work with both.
So let me save you the three weeks I lost figuring this out.
What each tool actually does
Veo 3 generates a scene. HeyGen generates a person talking.
Read that again.
Veo takes a text prompt and produces a short cinematic clip with sound: a drone shot over a job site, a product sitting on a counter with morning light, an animated character walking through a door. Nobody has to be on camera. There is no script being read to you. It is footage.
HeyGen takes a script and produces a human presenter delivering it to camera. Same face, same voice, every time, in as many languages as you need. It is a person, not a scene.
Different tools. Different jobs.
When I was the first salesperson through the door at a startup and we scaled that thing to nine cities, 136 internal reps and over 70 contractors across six countries, video outreach was the thing every rep wanted and nobody could sustain. Reps would film personalised intro videos for a week, then stop. Not because it did not work. Because filming forty of them on a Tuesday is miserable. That problem is a HeyGen problem. It was never a Veo problem. Knowing which bucket your problem falls into is the whole game.
What Veo is for
Veo is Google's video generation model, shipping as Veo 3.1 as of August 2026. You can reach it through the Gemini app, through Flow (Google's filmmaking interface built around it), and through the Gemini API and the Gemini Enterprise Agent Platform, which is the new name for Vertex AI since May 2026, if you want it wired into something programmatically.
The thing that changed everything for us was native audio. Veo generates sound with the video: ambient noise, sound effects, and dialogue with lip sync, all scored to the clip. You are not laying music over a silent generation anymore. You prompt the audio the same way you prompt the visuals.
Where it earns its keep for us:
- B-roll you cannot afford to shoot. Aerial shots, weather, locations, seasons. Things that would take a crew and a permit.
- Animated story content. Characters and worlds that do not exist. This is the bulk of what we generate.
- Ad visuals and product moments. Short, punchy, visually interesting clips built to stop a thumb.
Now the frustrating part.
Generations land in short clips, typically eight seconds at a time, now at up to 4K and in portrait as well as landscape. Veo 3.1 added scene extension and reference images, so you can feed it up to three stills of a character or product to hold the look across clips. That helps. It does not change the fundamental: longer pieces are stitched, not shot. Character consistency across clips takes real effort. Text rendered inside a video is unreliable. And you will burn a lot of credits on generations you throw away. Budget for waste, because it is not waste, it is the process.
What Veo output actually looks like
Talk is cheap, so here are real, unedited clips straight off our own production line. We generate personalized children's story videos under our Your Never Ending Story brand, and every second below came out of a text prompt. No camera, no crew, no animator.
And a completely different scene, same tool, same day:

And here is HeyGen, same week, same production
This is the comparison that actually matters, so here is ours rather than a vendor reel.
Same studio, same characters, same month. The clips above are Veo generating a whole scene.
This is HeyGen taking a
single still image of a character and making it talk:
Watch them back to back and the difference stops being theoretical. Veo built a world and
moved a camera through it. HeyGen took a picture and gave it a mouth. Neither tool could have
produced the other one.
Look at the bottom corner of our HeyGen clip. That is a watermark, and it is there because
of the plan the video was made on. Vendor demos never show you this. If these videos are going
in front of customers, check which tier removes the watermark before you build a workflow
around the cheap one.
If you need a specific person saying a specific thing, Veo is the wrong door entirely.
What HeyGen is for
HeyGen puts a presenter on screen. You can use their stock avatars or clone yourself, then feed it a script and get a talking-head video back. Their current Avatar V model builds a studio-grade, reusable presenter from about fifteen seconds of recorded footage. Worth knowing before you pick one: there are three kinds and they are not interchangeable. Stock avatars come ready made from HeyGen's library. Photo avatars are built from one still image, which is cheap and fine for volume. Digital Twins are trained from real video of a real person, and that is the one that actually looks like you.
Its real strength is volume and languages. Video translation handles a very large number of languages, and it clones the original voice so the translated version still sounds like the person, not a stranger dubbing over them. Its current avatar models handle lip sync, gestures and micro-expressions well enough that most viewers do not clock it in a short clip.
Where it earns its keep for us:
- Training libraries. Update a script, regenerate the module. No reshoot, no lighting setup, no matching your haircut from four months ago.
- Personalised outreach at volume. One script structure, many versions.
- Multi-language versions of one message. This is the killer feature and it is not close.
The frustrating part: avatars are still avatars. In a thirty second clip nobody notices. In a five minute clip, viewers start to feel that something is slightly off, even when they cannot name it. Energy is flat unless you write for it, because the avatar reads what you wrote, and if what you wrote is boring, you get a very well lit boring video. Voice cloning needs clean source audio. And credits move fast on the higher fidelity avatar models.
One more thing worth knowing: HeyGen has quietly become more than an avatar tool. Its Video Agent can draft a whole video from a description, and its editor now pulls in generated b-roll scenes using outside video models, including Veo itself. The lanes are blurring. The core strength has not moved though: a consistent presenter, at volume, in any language.
HeyGen does not make you interesting. It makes you repeatable.
Google Vids does talking heads too
Since mid-2026, Google has its own presenter product. Google Vids ships AI avatars inside Workspace, including a personal avatar built from a selfie and a voice recording, included with Workspace business plans and the Google AI Pro and Ultra subscriptions. The avatar library is younger than HeyGen's and the tooling is simpler, but the price is effectively zero if your team already lives in Google Workspace. Test it before you buy another subscription. If it covers your use case, you just saved a line item.
Both of these are the real procedures, verified against the vendors' own documentation this
month, written for an owner rather than an engineer. They include the pricing, the limits and
the parts that go wrong.
Download: Getting Veo Running Through Google Cloud
Download: HeyGen Setup Guide, API and MCP
Screens move. Both carry a "last verified" date on page one, so you always know how old the
steps are.
Veo 3.1, and what changed since Veo 3
Naming trips people up here, so this is worth ninety seconds. Google retired the Veo 3 model
IDs on 30 June 2026. That was a developer deprecation, not a rebrand, which is why almost
everyone still says Veo 3 and why you will keep seeing it written that way. The model you
actually generate with today is Veo 3.1.
What the newer version changed, in practical terms: you can feed it up to three reference
images so a character or a product holds its look across separate clips, it will extend an
existing scene rather than starting fresh every time, and it outputs in portrait as well as
widescreen at higher resolution. The eight-second clip length did not change, so longer pieces
are still stitched rather than shot.
If you are reading an older guide that tells you to call a Veo 3 model ID, that guide will
fail now. The rest of its advice is probably still fine.
Keeping the same face in every shot
Here is the problem that kills most first AI video projects. Every clip generates
separately, so your character drifts. The face shifts. The jacket changes colour. Your
protagonist ages four years between shot two and shot three.
The fix is a locked character block you paste into every single prompt, word for word. This
prompt writes that block for you. It follows the same five-part structure we use for every
prompt at ATP AI, Sales and Marketing: role, context, task, format, and questions before it
starts work.
ROLE: You are a video prompt engineer who specialises in keeping an AI-generated character looking identical across separately generated clips. CONTEXT: I am making short videos with an AI video generator. Every clip is generated on its own, so my character drifts between shots: the face changes, the clothing changes, sometimes the apparent age changes. I need one locked description I can paste into every prompt so the same character comes back every time. My character is: [DESCRIBE YOUR CHARACTER IN ONE SENTENCE]. TASK: Write me a reusable CHARACTER LOCK block, then show me how to use it across three different shots. FORMAT: 1. A CHARACTER LOCK block of 60 to 90 words covering face, hair, build, apparent age, clothing, and one distinguishing detail. Use concrete nouns. Avoid adjectives that could be read two ways. 2. Three example shot prompts. Each one opens with the CHARACTER LOCK block word for word, then describes only the shot. 3. A short list of the words that most often cause a character to drift, so I know what to avoid. 4. Give me all three parts as ONE complete markdown file, in a single code block, ready for me to save as character-lock.md and load into a Claude Project. Ask me up to 5 questions before you start.
Paste that into Claude and it will ask you up to five questions before it writes anything. Answer them properly, in your own words. Then it hands you back a finished file built around your character, not around our example.
That file is the thing worth keeping. Save it as character-lock.md. The .md on the end just means markdown, which is a plain text file with a little structure, and it is the format AI tools read best. If that is new, we cover it in what an MD file actually is.
Now open a Claude Project and add that file under project instructions, or upload it as project knowledge. Every chat you start in that project begins already knowing it. You stop re-explaining yourself every morning, which is the entire point.
Answer its questions with real detail and keep the block it gives you somewhere you can
paste from. That one habit is the difference between a set of clips that cut together and a
set that does not. It is the same discipline as the context files behind
a Claude Project:
write the thing down once, reuse it exactly, stop re-explaining yourself.
Which tool does which job
Here is the part to screenshot. Find your use case, take the tool.
| Use case | Right tool | Why |
|---|---|---|
| Sales outreach video to a named prospect | HeyGen | A human face saying the prospect's situation out loud. A generated scene cannot do trust. |
| Internal training library | HeyGen | Scripts change. Regenerating a module beats reshooting a person forever. |
| Social ad that has to stop the scroll | Veo 3 | Visual interest wins the first second. A talking head does not. |
| Product launch teaser | Veo 3 | Mood, motion, sound design. You are selling a feeling before a feature. |
| Multi-language customer onboarding | HeyGen | Voice-cloned translation across languages from one recorded source. |
| Trade show booth loop | Veo 3 | Silent-friendly, visually loud, plays on repeat without annoying anyone. |
| Founder explaining the offer | Neither, film it | Your real face for ninety seconds still outperforms both. Use the tools for what does not scale. |
Notice the last row. That one is free and most people skip it.
What it costs to start
Both run on credit systems, both change their pricing regularly, and both vary by region and billing cycle. So take structure from me and take numbers from them.
Veo 3 comes bundled into Google's consumer AI subscription tiers, which include a free tier with a limited monthly credit allowance, a mid tier that most small businesses land on, and a high-volume tier priced considerably higher. Developers can also pay per output second through the Gemini API or Vertex AI, which is the route to take if you are generating at real volume or automating it. Check Google's current pricing page before you commit.
HeyGen runs a free tier for testing, then paid plans that scale up through Creator, Pro and Business levels. Credits are consumed per minute of video, and the rate depends on which avatar model you use, with the higher fidelity models costing meaningfully more per minute than the older ones. There is also a pay as you go API. Check HeyGen's pricing page for current numbers.
Here is the honest framing: for most small businesses, one paid seat on each is a rounding error next to what a single day of production filming costs. The money is not the risk. The risk is buying both, using neither, and cancelling in March.
The mistake most owners make with AI video
They buy the tool before they have the job.
I have watched this pattern repeat with every client I work with. Somebody sees a Veo clip on LinkedIn, signs up, generates twelve gorgeous eight second clips of nothing in particular, and then has twelve clips of nothing in particular. There was never a piece of content those clips were for.
The fix is boring. Start with the asset you actually need. Write down the one video that, if it existed today, would make you money. Then pick the tool that makes that video. Not the other way around.
The second mistake is treating AI video as a content problem when it is usually a systems problem. A personalised outreach video is worthless if nothing routes it to the right prospect at the right stage. That is why we treat video as one piece of a larger AI integration build rather than a standalone toy. The video is the easy part. The delivery is the part that pays.
FAQ
Can you use Veo 3 for free?
Yes, up to a point. Google Flow gives you 50 free credits a day with no subscription, which is five Veo 3.1 Lite generations or two Fast ones. Not enough for production, plenty to find out whether it can do your job. HeyGen also includes some Veo credits in every account. Neither free tier gives you commercial rights without a watermark.
Does HeyGen have Veo 3.1 built in?
Yes. HeyGen ships Veo 3.1 as a selectable model in its editor, so you can generate a shot and chain it to your avatar in the same project without a separate Google account. That is why treating these as competitors is out of date. If the video is pure atmosphere with no presenter, going to Veo directly still gives you more control.
Do I need both?
Most businesses genuinely need one. We run both because we produce animated story content and presenter content as separate product lines. Start with one, prove it, then add the second if a real job demands it.
Will customers know it is AI?
Some will. That matters less than you think when the video is useful, and a lot when the video pretends to be something it is not. Do not fake a live testimonial. Do not clone a customer's face. Use it for training, explanation, translation and visuals.
Can you use Veo 3 for commercial use?
Not on the free tier. Commercial rights without a watermark sit behind a paid plan on both Google and HeyGen, and the tier you need depends on resolution and clip length. Check the current pricing page before you put anything in front of a customer, because both vendors move these numbers.
What to do this week
- Write the one video. One sentence: the video that would make you money if it existed on Friday. Not a content calendar. One video.
- Take it to the table above and pick the tool that job points to. Sign up for that one only. Ignore the other.
- Ship one real asset in seven days. Not a test. Something a prospect or a staff member actually watches. Send it, then look at whether anything happened.
If you want the systems view first, our work on AI for teams covers internal training and enablement, and AI for customers covers the front-facing side where most of this video ends up. More on my background and how I got here is on the founder page.
Still not sure which tool your use case needs? Book a call and bring your one video. We will tell you which tool builds it, or tell you to just film it yourself.
How to Install Claude Desktop - the four-tool setup, every screen answered
What Is an MD File? - the one-page files that brief your AI
What Is a Claude Project? - where your context files live and work
Do You Need Obsidian? - probably not, and here is the cheaper fix
Kaylee, the AI receptionist - the AI that answers when you cannot
Want more guides like this in your Google results?