Skip to content

An agent that sounds like your brand, not like a stock voice.

Custom Voice Cloning

We train a bespoke synthetic voice from recordings you supply, with documented speaker consent, and deliver it restricted to your tenant. One voice, one-off fee.

  • One bespoke voice, locked to your tenant
  • Documented speaker consent required
  • Full SSML control
  • Revocable by you at any time
Choose a plan

Licence keys are issued the moment payment clears. Prices exclude VAT, which is calculated at checkout from your billing country.

Overview

A preset voice is fine until the agent becomes the primary way customers hear from you. At that point the voice is brand, and a stock voice used by a hundred other companies stops being acceptable.

Custom Voice Cloning produces one bespoke voice trained on recordings you provide. You need roughly thirty minutes of clean speech from a single speaker, along with that speaker's written consent, which we record and retain against the voice. We handle the training, the quality review and the tuning pass, and deliver the voice locked to your tenant.

The finished voice supports the same SSML controls as our presets — pace, emphasis, pauses — and can be revoked by you at any time, which permanently disables it.

Consent is not a formality here. We will not train a voice from recordings of a person who has not given documented, informed permission, and we will not clone a public figure or a third party's voice talent without evidence of the right to do so.

Capabilities

  • Trained on your speaker

    Around thirty minutes of clean single-speaker audio produces a voice that carries the speaker's timbre and cadence rather than an approximation.

  • Consent on record

    The speaker's written consent is recorded against the voice and retained for the life of the voice, which is what makes this defensible rather than merely possible.

  • Full prosody control

    Pace, emphasis, pauses and pronunciation overrides work exactly as they do with preset voices, using the same SSML.

  • Tenant-locked

    The voice is usable only within your tenant. It is never added to a shared voice library or made available to another customer.

  • Revocable

    You can revoke the voice at any time, which permanently disables synthesis with it. Revocation is recorded in the audit log.

  • Human quality review

    Every voice goes through a listening pass against a standard script before delivery, with one tuning iteration included.

Where it earns its keep

  • Consistent brand voice

    Use the same voice across your phone agent, IVR prompts, in-product narration and video, so customers hear one identity.

  • Named spokesperson

    Extend the reach of a founder or presenter whose voice customers already recognise, with their explicit consent.

Questions

What happens if I cannot provide speaker consent?
We will not train the voice. This is a firm limit, not a negotiable one — it protects the speaker, you, and us. If the speaker is available but the paperwork is not, we can provide a consent template.
Can you clone a celebrity or a voice actor's voice?
Only with documented permission from that person or their representative covering the intended commercial use. Without it, no.
What if the delivered voice is not good enough?
One tuning iteration is included. If after that iteration the voice still fails the agreed quality bar, we refund the fee in full.
How long does it take?
Ten business days from the point we accept your audio. Audio that needs re-recording because of background noise or inconsistent levels is the usual cause of delay, so we review it before the clock starts.
Is the voice used to improve your shared models?
No. The training data and the resulting voice are confined to your tenant and are deleted on request.