Videos · Glossary

What is AI voice cloning?

AI voice cloning is training or conditioning a speech model on recordings of one specific person so it can say new words in that person's voice. The clone speaks any text you give it with their tone, accent and delivery. Some current systems need only seconds of reference audio, which is why consent now matters so much.

A cloned voice can dub a founder's video into five languages, and the same technique can fake a phone call from a relative asking for money. The technology is identical in both cases, so the rules focus on consent and on what the voice is used to say.

· Co-founder

5 min read · Published

Voice cloning and stock text to speech compared
Voice cloningStock TTS voice
Whose voiceA specific real personA voice built by the provider for general use
What it needsRecordings of that person; Google's Instant Custom Voice takes up to 10 seconds plus a recorded consent statementOnly text
ConsentNeeded from the person whose voice it isHandled by the provider's own licensing
AccessOften restricted; Google limits Instant Custom Voice to allow-listed usersGenerally available
Main legal exposureImpersonation, fraud and misleading endorsementPlatform and advertising disclosure rules
YouTube disclosureMaking a real person appear to say something they did not must be disclosed; cloning your own voice for voiceovers or dubs need not beNot a depiction of a real person
In the video builder hereNot offeredThe only kind of voice used

One person's voice, any words

Ordinary text to speech speaks in voices a provider designed and offers to every customer. A clone is personal. It captures what makes one voice recognisable, such as pitch, timbre, accent and rhythm, and applies it to sentences that person never recorded. Early versions needed long studio sessions and still sounded mechanical. Current systems can produce a convincing likeness from a short clip, and that change in effort is the whole story: when a few seconds of audio from a public video is enough, anyone who speaks on camera has, in effect, published the raw material for a copy.

How a clone is made

The provider collects reference audio, analyses it, and conditions a synthesis model so that new speech matches it. Responsible services pair that with proof of permission. Google's Chirp 3 Instant Custom Voice is a clear example of the pattern. It takes two single channel recordings of up to 10 seconds each: one of reference speech with no background noise, and one of the speaker reading a fixed consent statement, in English, I am the owner of this voice and I consent to Google using this voice to create a synthetic voice model. A custom consent script is not accepted, and access is restricted to allow-listed users who request it through Google's sales team.

Consent, fraud and the law

Regulators treat cloning as a fraud risk first. The Federal Trade Commission in the United States ran a Voice Cloning Challenge on the harms, including fraud and extortion scams, and in April 2024 pointed to its new Impersonation Rule as a tool against deceptive cloning, stating that there is no AI exemption from existing law. In Europe, Article 50 of the AI Act applies from 2 August 2026: deployers must disclose deep fakes, and providers must mark synthetic audio in a machine readable way. Australia's eSafety Commissioner publishes guidance for industry on deepfakes, a category that covers faked voices as well as faces. None of this bans cloning; it targets deception and use without permission.

When cloning is legitimately useful

The strongest cases involve a person using their own voice. A founder can dub product videos into other languages without losing the voice customers already know. A course creator can fix a mistake in a lesson by editing text instead of booking a studio. YouTube's disclosure policy draws the line in exactly this place: cloning your own voice to create voiceovers or dubs is listed among uses that need no label, while making a real person appear to say something they did not must be disclosed. Using anyone else's voice needs their clear, written permission and a defined scope.

Common mistakes

Cloning a colleague, presenter or former employee on the strength of a verbal yes is the first, because consent needs to be recorded and to name the uses allowed. Assuming one consent covers everything is the second; permission for training videos is not permission for advertising. Putting a cloned voice behind a customer testimonial is the third, since an endorsement must reflect a real person's honest experience. Storing voice samples casually is the fourth, as they are biometric material. And skipping disclosure where a platform or law requires it risks removal of the video and the account.

Where it shows up in the product

The video builder here does not clone voices. There is no upload of a voice sample, no custom voice training and no way to make narration sound like you or anyone on your team. Every line is spoken by a Google Chirp3-HD stock voice: the generator picks from four styles in two genders, energetic, instructional, warm and neutral, and the voice dialog lets you search the whole live Google catalogue per scene. In a UGC video the creator persona is also synthetic, one generated portrait and one stock voice speaking every line, not a copy of a real person. If a video must carry a specific person's voice, record that person.

Questions people ask

Is voice cloning legal?

The technique itself is generally legal. What creates legal exposure is how it is used: copying someone's voice without permission, using it to deceive, impersonating a business or official, or putting words in a real person's mouth in an advert. Rules differ by country and are changing quickly, so written consent and clear disclosure are the minimum for any commercial use.

How much audio does a voice clone need?

It varies by provider and quality level. Some professional services still ask for long, clean recordings to build a high quality custom voice, while instant services work from very short samples. Google's Instant Custom Voice, for example, takes a reference recording of up to 10 seconds, alongside a separate recording of the speaker reading a consent statement.

Can a cloned voice speak other languages?

Some systems can apply a voice to languages the speaker never recorded, which is what makes dubbing with your own voice possible. Quality varies by language pair, and accents and pronunciation can drift from what a native speaker would expect. Anyone using it for customer facing content should have a fluent speaker review the result before it is published.

How can a business guard against cloned voice scams?

Treat a voice alone as unproven for any request involving money, passwords or account changes. Call the person back on a number you already hold, agree a verification step for urgent payment requests, and train staff that a familiar voice on the phone is no longer proof of identity. Report attempted scams to the relevant consumer or fraud authority.

Does using a stock AI voice avoid the consent issue?

Largely, yes, because a stock voice is not modelled on a particular person you need permission from; the provider handles the licensing of its voices. Advertising rules still apply to what the voice says. A stock voice presented as a real named customer giving a testimonial is still a misleading endorsement, even though nobody's voice was copied.

Make one with videos

The button opens the generator with this use case already described. Change the wording to match your own.

Create a video with OneCraft

Related questions

Step by step in the builder: Add voiceover, captions and music to a video.

Sources

Written and checked by the OneCraft team. Last checked .