PROD. SECOND UNIT UGC/SIDES — MODULE 2/DRAFT/Fact-check before publish

Living referenceModule 2  ·  section 5  ·  3,724 words · ~14 min read

Tool landscape — voice, avatar, and editing

living-referenceDoc type
draftStatus
RequiredFact-check
2026-08-14Last verified
Document controlLiving reference — decays by design
Last verified
2026-08-14
Method
Tool list: web-search snapshot 2026-08-14. PRICING: vendor pricing pages fetched directly 2026-08-14 for HeyGen, Synthesia, Creatify, Colossyan, Captions; Arcads publishes none; VEED's page did not render to this method. No tool tested. Q1 (commercial licence) and Q4 (data retention) still unverified.
Cadence
90 days maximum, and immediately on any acquisition, shutdown or pricing-model change in a named tool
Owner
TBD — needs a named owner before publish (blueprint §3.2 Caution pattern)no named owner
Read this before anything below it

This is a snapshot search, not a verified market survey.

The tools named here were assembled from a single web-search pass on 2026-08-14. Pricing was subsequently verified against vendor pricing pages on the same date and is the only part of this document that has been. No tool was tested by anyone who wrote this. Capability claims, commercial licence terms (Q1) and data retention (Q4) remain unchecked.

Confirm pricing, features, licence terms and availability against each tool's own site before you act on anything in this document — and particularly before you tell a client what a tool can do.

This page has a last_verified date in its front matter and a 90-day maximum cadence. If that date is more than 90 days old, treat this page as a list of things to go and check, not as information.

1Why this document is organised the way it is

03-voice-and-avatar-landscape.md teaches the category and seven evaluation questions. This document is where those questions meet actual product names. It does not rank, and the refusal to rank is deliberate — see below.

The seven questions, in short form, since they're referenced throughout:

Question
Q1 Can I use the output commercially, and can I prove it?
Q2 Whose voice or face is this, and what did they agree to?
Q3 Will I be able to get this exact voice back in six months?
Q4 What happens to the sample I upload?
Q5 Can I hear the tell — and can my audience?
Q6 Does it handle the way people actually speak in my market?
Q7 What does it cost per finished minute, not per month?

2A finding about the sources, which changes how you should read every list like this

The search behind this page turned up perhaps thirty "best AI UGC tool" pages. Almost none of them were independent. The pattern, in rough order of frequency:

  • Vendor blogs ranking their own category. A tool's own site publishing "7 best AI caption generators" or "11 best alternatives to [competitor]".
  • Vendor blogs ranking themselves first. At least two pages placed the publishing company's own product at number one in a list it had written.
  • Affiliate review sites with a consistent house style and no disclosed testing methodology.
  • Directory and comparison sites whose business is being the comparison.

One page ranked a tool this document has excluded at #1 with no corroboration anywhere else in the search. That is the shape to watch for, and it generalises: a tool appearing at the top of exactly one list, and nowhere else, is a marketing placement until proven otherwise.

This is why the document below is organised by category and by evaluation question rather than as a ranking. A ranking would imply a confidence that nothing available to produce this page could support.

3How confidence is marked

Marker Means
Corroborated Named consistently across several independent-ish sources this pass
Named only Appeared, but with no detail this document is willing to restate
Unverified detail A specific capability claim seen once or twice and not confirmed
No data Explicitly not captured — usually pricing

Pricing was filled in on 2026-08-14 from vendor pricing pages, replacing the No data that stood here previously. See the pricing table below. Five of seven published usable figures; one publishes none; one did not render.

Q1 (commercial licence terms) and Q4 (data retention) are still No data across the board. Those are the two questions a client's legal team asks, and this page still cannot answer either.

The pricing pass covered seven tools only — the avatar/spokesperson category plus VEED and Captions. The voice tools (ElevenLabs, Resemble AI, PlayHT) and the remaining editors (Descript, CapCut, Opus Clip) still show No data for Q7 and are the obvious next batch.


4Pricing — verified 2026-08-14

Fetched from each vendor's own pricing page on the date above. Monthly-billing figures unless stated. Every one of these will move; the whole point of the last_verified date is that you check rather than trust.

Tool Free tier Entry paid Mid Top published Enterprise
HeyGen Yes — 3 videos/mo, ≤1 min Creator $29 ($24 annual) Pro $49 ($40) Business $149 ($122) Contact sales
Synthesia Yes — 10 min/mo, 9 avatars Starter $29 ($18 annual) Creator $89 ($64) Contact sales
Creatify Yes — 10 credits/mo, watermarked Starter $39 Pro $99 Contact sales
Colossyan Yes — 20 min/mo Professional $59 ($30/seat annual) Contact sales
Captions Yes — no generative credits Max $24.99 Scale 1x $69.99 Scale 4x $279.99 Contact sales
Arcads Trial promo only Not published
VEED Yes Not captured

The two gaps, stated rather than left blank:

  • Arcads publishes no pricing. The public site shows a promotional trial and routes you to log in for rates. That is a real finding — you cannot compare its cost against the others without giving them your details, which is itself worth knowing before you start evaluating.
  • VEED's pricing page did not render to the method used here. Third-party sources disagree substantially — figures seen for the same tier ranged roughly $9–$24 for its entry plan and $24–$35 for the next one. Those are not asserted here. Open the page yourself; do not take a comparison site's number.

Three things the table shows that a range doesn't

1. Real free tiers are the norm, not the exception. Five of the five tools with published pricing have one — 3 videos, 10 minutes, 20 minutes, watermarked credits, editing-without-generation. They are limited, and they are enough to evaluate a tool properly before paying. Nobody has to spend anything to start.

2. Annual billing moves the floor a long way. Synthesia Starter drops from $29 to $18; HeyGen Creator from $29 to $24; Colossyan Professional from $59 to $30 per seat. If you commit for a year the entry point is meaningfully below the monthly sticker.

3. The top end goes higher than a single subscription suggests. Captions Scale 4x is $279.99/month on its own. And a working setup is usually two tools, not one — an avatar or voice tool plus an editor. Two entry tiers is around $50–70/month; two mid tiers is $150–250.

4. Nothing here is unlimited. Every tool with published pricing meters — credits (HeyGen, Creatify, Captions) or minutes (Synthesia, Colossyan). Synthesia Starter is 10 minutes a month. HeyGen Creator is 600 credits. Colossyan Professional is 30 minutes. These are subscriptions with allowances, not all-you-can-eat. This directly contradicts language elsewhere in the product — see register A6.

What this does not tell you

Credit systems make the sticker price a poor guide to cost per finished video. HeyGen, Creatify and Captions all meter in credits; Synthesia and Colossyan meter in minutes. Q7 asks what it costs per finished minute, and none of these pages answer that — you only learn it by running your own work through and dividing. That remains the honest answer to Q7, and no pricing table replaces it.


5Category 1 — Avatar and spokesperson tools

A synthetic person delivers your script to camera. Either a stock presenter from the vendor's library or a likeness built from footage of a real person.

Category-level caution — applies to every tool below, without exception

Q2 is a category property, not a product feature. No tool in this category can give you permission you don't have.

  • A stock presenter from a vendor's library is licensed by the vendor. Q1 becomes "which tier permits commercial use, and where is that in writing."
  • A likeness of a real person — you, a client's founder, anyone — is a consent question first. Written permission naming the specific use is the floor. Search results this pass consistently indicated that a general terms-of-service checkbox is not treated as sufficient for commercial use of another person's likeness or voice, and that unauthorised commercial use creates real exposure. That is a legal question and this document is not legal advice — see register item C1.
  • Q5 in this category is a whole-body problem, not a voice problem. Gaze, blink rate, micro-gesture and the join between head and torso all give it away before the audio does.

HeyGen — Corroborated

Avatar video generation with a heavy emphasis on multilingual output. Sources this pass repeatedly associated it with very wide language coverage (one figure seen: 175+ languages — unverified detail).

  • Q6 — likely its strongest answer. If your market is not English-first, this is the category entry that came up most often for that reason.
  • Q1, Q2, Q4 — not captured. Ask directly.
  • Q7 — free tier, then $29/mo ($24 annual), up to $149 Business. See table.

Synthesia — Corroborated

The enterprise-positioned entry. Consistently described as avatar-plus-language at scale, aimed at corporate video rather than social ads specifically.

  • Q1 — enterprise positioning usually correlates with clearer commercial terms, and usually correlates is not verified. Check.
  • Q6 — wide language coverage claimed (140+ languages, 160+ avatars seen — unverified detail).
  • Q5 — the enterprise register cuts both ways for UGC: polish is the opposite of what most UGC briefs want.
  • Q7 — free 10 min/mo, then $29/mo ($18 annual — the lowest entry point found). Metered in minutes rather than credits, which makes cost-per-video easier to predict than most.

Arcads — Corroborated

Positioned specifically at UGC-style ad creative rather than corporate video, with a library of actor-style avatars (200+ seen — unverified detail).

  • Q5 — the one in this list most explicitly aimed at not reading as corporate. Whether it succeeds is a judgement you have to make with your own ears and eyes, on your own script.
  • Q2 — a library of "actors" raises the obvious question of what those performers agreed to and for how long. Ask. Get the answer in writing.
  • Q1, Q4 — not captured.
  • Q7Arcads publishes no pricing. Log-in required. The only tool here you cannot price without handing over contact details.

Creatify — Corroborated

Associated with high-volume batch generation for e-commerce, including URL-to-video style workflows.

  • Q7 — free tier (10 credits, watermarked), Starter $39, Pro $99. Credit- metered, and this is the entry where that matters most: batch tools are where cost per finished video diverges hardest from the monthly sticker. The sticker is now known; the per-video number still is not.
  • Q5 — volume and tell are in direct tension. Twenty near-identical videos read as twenty near-identical videos.

Colossyan — Corroborated

Script-to-video with avatars, consistently described as strongest for training, onboarding and internal communications rather than social ads.

  • Fit is the question here, not capability. Several comparisons this pass drew the same line: Colossyan for structured/instructional, something else for social-first. If a brief is UGC, this may be the wrong shelf.
  • Q6 — multilingual and lip-sync capability referenced. Unverified detail.
  • Q7 — free 20 min/mo, Professional $59/mo, or $30 per seat annually — the largest monthly-to-annual gap in the table.

Captions — Corroborated (also appears in Category 3)

Mobile-first, and the entry that most blurs the avatar/editing line. Features referenced this pass include a personal digital twin, gaze correction, and one-tap audio cleanup (unverified detail).

  • Q2 — read this one carefully. A digital twin of yourself is the cleanest consent case in the category. It is also the feature most likely to be used casually, without the written record you'd want if a client later asks who authorised it.
  • Q5 — gaze correction is a tell-management feature. It is also, arguably, a tell of its own once you know to look for it.
  • Q7 — free tier does editing but no generative credits; Max $24.99, then $69.99 / $139.99 / $279.99 by credit volume. Note the vendor states these are iOS plans, so confirm the price on the platform you'll actually buy on.

6Category 2 — Voice cloning and synthetic voice

Text in, speech out — from a stock voice, or from a clone of a specific person.

Category-level caution — Q2 and Q3, and Q3 is the one people forget

Q2: consent is not a feature any vendor can supply. Every source this pass that addressed the legal side landed in the same place: cloning another person's voice for commercial use requires explicit, documented permission naming the use, and a terms-of-service checkbox is not that. Register item C1 covers this; it needs a lawyer, not a comparison table.

Q3 — and here is a real, dated example rather than a hypothetical:

PlayHT / PlayAI was acquired by Meta in July 2025. This is corroborated by TechCrunch and by a law-firm transaction notice. At least one directory source further stated the service was subsequently shut down — that part is single-source and this document does not assert it.

The lesson survives the uncertainty. If you had built a client's voice on that platform, the question "can I get this exact voice back in six months" stopped being theoretical on a specific date, and no feature comparison would have warned you. Ask what happens to your clone if the company is acquired, if you stop paying, or if the voice is retired. Get it in writing. Keep exports where the format allows.

ElevenLabs — Corroborated

The name that appeared most often, and the one most often used as the benchmark others compare themselves against.

  • Q5 — most sources this pass put it at or near the top for naturalness and emotional range. One cited a word-error-rate figure of 2.83% (unverified detail, and WER measures transcription accuracy, not whether a line sounds human — do not confuse the two).
  • Q4 — not captured. This is the question to ask before uploading a client's founder's voice.
  • Q7No data.

Resemble AI — Corroborated

Consistently described as API-first, with fine-grained control and — notably — watermarking and provenance features.

  • Q2 and Q4 — the strongest positioning in this category on paper. Ethical sourcing and watermarking came up repeatedly. Whether the implementation meets the claim is not something this document has checked.
  • Q5 — emotional parameter control referenced. Unverified detail.
  • If you are cloning a client's voice rather than your own, this is the category entry whose Q2/Q4 answers are worth reading first.

PlayHT / PlayAI — Corroborated, with a status warning

See the Q3 example above. Historically positioned as lower-cost with a larger voice library and low-latency API access.

  • Status is the finding. Acquired by Meta, July 2025. Availability and terms after that point are unverified and possibly moot.
  • Do not start a client series on this without confirming it is still operating and what your exit looks like.

Typecast, Murf, Cartesia, Camb.ai — Named only

All appeared in comparison material this pass. None had enough corroborated, independent detail to justify a description here. Listed so a future verification pass knows where to look, not as recommendations.


7Category 3 — Editing and assembly

The passes discussed in 04-assembly-and-editing.md: transcript-driven cutting, filler removal, captions, reframing, clipping.

Category-level caution — the questions shift

Q2 barely applies — you're editing your own footage, and the consent questions that dominate categories 1 and 2 mostly aren't live here.

Q4 becomes the important one instead. You are uploading a client's unreleased footage to a third party. What happens to it, how long is it retained, is it used for training, and can you delete it? Almost nobody asks, and it is the question a client's legal team will ask you.

Q5 changes shape too. The tell here isn't synthetic delivery — it's over-processing. Machine-perfect filler removal produces a cadence no person has, and burned-in auto-captions with a default animation are their own genre signal.

Descript — Corroborated

The transcript-driven editor: edit the words, the video follows. Repeatedly named as the strongest entry for talking-head and interview material.

  • Maps directly onto §4's "mechanical and reversible" test — this is the archetype of a pass worth handing over.
  • Q4 — not captured. Ask before uploading client footage.

CapCut — Corroborated

Broad free-tier social editor: captions, auto-reframe, background removal, templates. Consistently described as the default for TikTok/Reels/Shorts-shaped work, with a free tier and no export watermark (unverified detail — verify, as watermark and free-tier policies change often and are exactly the sort of thing that changes quietly).

  • Q1 — free tiers and commercial use are where licence terms most often surprise people. If you are billing a client for output made on a free tier, read that tier's terms.
  • Q7 — free is not zero if it costs you a rights problem.

Opus Clip — Corroborated

Long-form to short-form clipping: upload a long video, get multiple clips with captions and platform framing.

  • Fits §4's "handed over" column cleanly, with one caveat: clip selection is a taste judgement dressed as a mechanical one. It picks what to keep. Watch the output before trusting the selection, not just the captions.

VEED — Corroborated

Browser-based timeline editor with AI captioning, subtitling, translation and resizing. Positioned as desktop-first and closer to a conventional editor than the mobile-first entries.

  • Sits between "editing" and "assembly" — useful where you want the AI passes but still want a timeline underneath them.
  • Q4 — not captured.
  • Q7not captured. The pricing page did not render to the method used here and third-party figures conflict. Open it yourself.

Captions — Corroborated (see Category 1)

Mobile-first, strongest on rapid captioning for short-form, plus the avatar features noted above.

Submagic, Runway — Named only

Submagic appeared in captioning comparisons; Runway appeared as generative video rather than editing. Neither had enough corroborated detail this pass. Listed for the next verification pass.


8What was excluded, and why

Two tools were named at or near the top of exactly one source each, with no corroboration anywhere else in the search, on pages published by parties with a commercial interest in the ranking. They are not named here.

Naming them would give a marketing placement the same standing as a tool that appeared across a dozen sources, and this document's only real value is the distinction between those two things.

The rule going forward, for whoever owns this page: a tool enters this list when it appears across multiple sources that do not sell it. One enthusiastic page is not a source. A tool's own blog is never a source about its own category.


9Re-verification protocol

For whoever picks this up — and per the front matter, that person does not exist yet and must before publish.

Every 90 days, or immediately on news of an acquisition, shutdown or pricing-model change:

  1. Confirm every named tool still exists and still does what this page says. Start with PlayHT/PlayAI, whose status is already flagged.
  2. Re-check pricing. Done once, 2026-08-14 — see the pricing table. Every figure in it is perishable, and two gaps remain open: Arcads publishes nothing, and VEED did not render. Re-run this whenever the date at the top passes 90 days.
  3. Fill in Q1 and Q4 for at least the tools in categories 1 and 2. Commercial licence terms and data retention are the two questions a client's legal team asks, and right now this page cannot answer either.
  4. Re-run the source-quality check. The vendor/affiliate dominance described above will not have improved. Re-check whether any genuinely independent testing has appeared.
  5. Promote or demote confidence markers based on what you find, and update last_verified.
  6. Delete anything that has gone quiet. A tool nobody has written about in six months is a tool a reader should not build a client workflow on.

If you cannot do a full pass, do step 1 only and update the date. A page that honestly says "these still exist, we haven't checked anything else since March" is far more useful than a page that silently implies everything on it is current.


10Sources

Snapshot search, 2026-08-14. Listed for the next verifier's benefit, including the low-quality ones, so the source-quality finding above can be re-checked rather than taken on trust.

Your word for it — nothing is tracked automatically.

Source: content/module-2/05-tool-landscape.md