A preset voice is fine until the agent becomes the primary way customers hear from you. At that point the voice is brand, and a stock voice used by a hundred other companies stops being acceptable.
Custom Voice Cloning produces one bespoke voice trained on recordings you provide. You need roughly thirty minutes of clean speech from a single speaker, along with that speaker's written consent, which we record and retain against the voice. We handle the training, the quality review and the tuning pass, and deliver the voice locked to your tenant.
The finished voice supports the same SSML controls as our presets — pace, emphasis, pauses — and can be revoked by you at any time, which permanently disables it.
Consent is not a formality here. We will not train a voice from recordings of a person who has not given documented, informed permission, and we will not clone a public figure or a third party's voice talent without evidence of the right to do so.
