Skip to main content
In the Voice section, you choose and customise exactly how your agent sounds to the customer. It has a few major parts.

Voice Provider

There are 4 main providers. The rupee figures below are rough cost guides.
Note: Avoid using emojis and special characters in the system prompt — they can break the agent’s behaviour during a live call.

Voice Models

Every provider has its own voice models. A voice model’s job is to turn text into speech (this is called TTS, or Text To Speech).

Voice

This is the actual voice the customer hears on the call. Every provider offers different voices, and each voice has its own character — some are sharp, energetic, and engaging, while others are calm, soothing, and friendly.

Voice Parameters

Each provider also gives you settings to fine-tune the voice:
  • Voice Speed (ElevenLabs, Cartesia, Sarvam, Azure): controls how fast or slow the agent talks.
  • Voice Stability (ElevenLabs): decides whether the agent sounds like a real human or a robot.
  • Voice Similarity Boost (ElevenLabs): how closely the agent’s voice matches the original voice it was trained on.
  • Voice Style (ElevenLabs): makes the voice more expressive and dramatic — like turning up the emotion dial.
  • Voice Volume (Cartesia, Azure): controls how loud or soft the agent speaks.
  • Voice Pitch (Sarvam, Azure): makes the voice younger and brighter, or deeper and more mature.