Avatar models + rendering
Voice cloning
Voice cloning is the process of creating a synthetic voice that resembles a specific speaker, so an avatar can speak with a chosen voice while still generating new responses in real time.
How voice cloning works
Voice cloning creates a synthetic voice that resembles a specific speaker. The goal is to preserve recognizable qualities such as tone, accent, pacing, and timbre while allowing the system to speak new text.
In an avatar system, voice cloning is usually part of the text-to-speech layer. The agent generates a response, TTS produces audio in the cloned or selected voice, and the avatar's face animates to match that speech.
A concrete example: a brand might use an approved voice for a product guide so every real-time avatar session sounds consistent across support, onboarding, and sales.
Voice cloning needs clear consent, ownership, and safety controls. A realistic voice can be powerful, so teams should define who can create, use, and modify cloned voices.
What Anam ships
Anam's Cara-4 model delivers expressive real-time avatars with around 150 ms server-side avatar-generation latency once a session is running, across 70+ languages. Builders use JavaScript and Python SDKs or integrations for LiveKit, Pipecat, ElevenLabs Agents, Agora, and VideoSDK. Bring any AI stack including OpenAI, Claude, Gemini, Mistral, Groq, Deepgram, Cartesia, or custom providers. The platform supports WebRTC delivery, SOC 2 Type II, HIPAA, zero data retention, and regional data residency. Sessions stream low-latency audio and video to browsers and native apps.
Related terms
Frequently asked questions
What is voice cloning in avatar systems?
Voice cloning creates a synthetic voice that resembles a chosen speaker, so an avatar can speak new generated responses using that voice profile.
Is voice cloning the same as text-to-speech?
Voice cloning is a way to create or select a voice identity. Text-to-speech is the system that turns text into spoken audio using that voice.
What permissions are needed for voice cloning?
Teams should have explicit consent from the speaker, clear usage rights, access controls, and policies for where the cloned voice may appear.
What makes a cloned voice work well with an avatar?
It should sound natural, start quickly, handle pronunciation reliably, match the avatar persona, and provide stable timing for lip sync during live sessions.
Last updated: 17th July 2026 · Reviewed quarterly.
Try the real-time avatar API trusted by 8,000 builders
© 2026 Anam Labs
HIPAA & SOC 2 Certified