TTS Generator
Convert text to natural-sounding speech using AI-powered text-to-speech synthesis. Generated audio is automatically saved to your Audio Library for immediate use in campaigns and IVR flows.
Overview
The TTS (Text-to-Speech) Generator transforms written text into spoken audio files. It supports multiple Indian languages and speaker voices, making it ideal for creating localized voice content without recording studio time, voice actors, or audio editing software.
Supported Languages and Speakers
Sarvam AI (Indian Languages)
| Speaker ID | Gender | Voice Characteristics |
|---|---|---|
| ishita | Female | Warm, natural cadence with a friendly tone suited for greetings. |
| shubh | Male | Professional, clear delivery ideal for business messages. |
| shreya | Female | Conversational, approachable tone for general messaging. |
| manan | Male | Steady, measured pace for informational content. |
| aayan | Male | Direct and authoritative, well-suited for notifications. |
Deepgram Aura (English)
| Speaker ID | Voice Characteristics |
|---|---|
| aura-helios-en | Bright, energetic voice for promotional content. |
| aura-asteria-en | Warm, inviting tone for greetings and welcome messages. |
| aura-luna-en | Soft, calm delivery for customer support and sensitive topics. |
| aura-stella-en | Clear, professional voice for business communications. |
| aura-athena-en | Authoritative, confident tone for formal announcements. |
| aura-zeus-en | Deep, resonant voice for impactful messaging. |
| aura-arcas-en | Natural, conversational style for general content. |
| aura-perseus-en | Polished, articulate delivery for IVR prompts. |
Speaker availability: The speaker list is subject to provider availability. Not all speakers may be available for all languages. Test speaker previews before generating production content.
Use Cases
Welcome Messages and Greetings
Generate personalized welcome messages for new users. Use TTS to dynamically insert names, company names, or other variables into voice prompts without recording individual files.
Example: “Hello [Customer Name], welcome to Acme Services. Your account has been activated successfully.”
Notifications and Alerts
Create timely audio notifications for transactional events — payment confirmations, order status updates, appointment reminders, or service alerts.
Example: “Your electricity bill payment of rupees two thousand five hundred was received on July 15th. Thank you.”
Multi-Language Campaigns
Build campaigns targeting audiences across different regions by generating the same message in multiple languages. Upload all variants to the Audio Library and segment your campaign contact lists by language preference.
Example workflow:
- Write your message in English.
- Translate and generate TTS variants in Hindi, Tamil, Telugu, and Bengali.
- Segment contacts by preferred language.
- Launch four parallel campaigns, each using the appropriate language audio file.
Rapid Prototyping for IVR Flows
During IVR flow development, use TTS to quickly generate placeholder prompts for testing. Once the flow logic is finalized, replace TTS-generated prompts with professionally recorded audio if desired.
Dynamic Content
For campaigns where message content changes frequently (e.g., daily deals, stock prices, weather alerts), TTS eliminates the need to re-record audio for every update.
Generating TTS
Step-by-Step
- Navigate to Audio Library from the sidebar.
- Click Generate Voice.
- Select the Target Language from the dropdown.
- Choose a Speaker Voice — preview a sample by clicking the play icon next to each speaker name.
- Enter your Text in the text area. Use the formatting tips below for best results.
- Click Generate Voice. Processing typically takes 3–10 seconds depending on text length.
- The generated audio file appears in your Audio Library with a name matching the first few words of your text. You can rename it for clarity.
Text Formatting Tips for Better TTS Output
The quality of TTS output depends significantly on how you structure your input text. Follow these guidelines for the best results:
- Use proper punctuation. Periods, commas, and question marks create natural pauses and intonation. A well-punctuated sentence sounds far more natural than a run-on block of text.
- Spell out acronyms and abbreviations. Write “United States” instead of “U.S.”, “Doctor” instead of “Dr.”, “January 15th” instead of “15/01”. The TTS engine interprets these phonetically and may mispronounce abbreviations.
- Write numbers as words for critical values. “Your balance is rupees two thousand five hundred” sounds clearer than “Your balance is rupees 2500.00” for financial amounts. For phone numbers, write digit by digit: “nine eight seven six five four three two one zero”.
- Use line breaks to create pauses. Each new line or paragraph adds a brief pause in the generated speech, helping to separate logical sections of your message.
- Keep sentences under 20 words. Longer sentences strain the TTS engine’s prosody model and can lead to unnatural rhythm.
- Avoid special characters. Symbols like @, #, $, %, &, *, and emojis are either mispronounced or produce garbled output. Write out the words they represent.
- Specify pronunciation for ambiguous words. For words with multiple pronunciations (e.g., “lead” as in metal vs. guidance, “read” as present vs. past tense), the engine picks the most common. Test and adjust phrasing if the output is incorrect.
Character limits and cost: Each TTS generation is limited to 1,000 characters per request. For longer messages, split the text across multiple generation requests and combine the audio files in your audio editor. Each TTS generation consumes a small amount of voice minutes (approximately 0.1 voice minutes per 100 characters), which is deducted from your voice minute balance. Plan your budget for large-scale TTS content creation.
Combining TTS with Audio Files in Campaigns
TTS-generated audio and pre-recorded audio files can be used together in a single campaign for a polished, professional result:
| Approach | How It Works |
|---|---|
| TTS intro + recorded body | Generate a personalized greeting with TTS (e.g., including the recipient’s name), then play a professionally recorded message body. |
| Recorded intro + TTS content | Use a recorded jingle or brand intro, followed by TTS-generated dynamic content (e.g., order details or appointment times). |
| TTS for all content | For rapidly changing content like daily specials, use TTS exclusively. Accept the slight quality trade-off for speed and flexibility. |
| IVR hybrid | In IVR flows, use recorded audio for static prompts (e.g., “Welcome to the main menu”) and TTS for dynamic information (e.g., “Your current balance is…”). |
Troubleshooting
Issue: TTS generation fails with an error.
Cause: API configuration issue — the TTS engine requires a valid AI provider API key.
Solution: Verify that your OpenAI or Gemini API key is configured and active in Settings . Test the key by generating a short TTS sample.
Issue: Generated audio sounds robotic or unnatural.
Cause: Poorly formatted input text — run-on sentences, missing punctuation, or excessive special characters.
Solution: Review the Text Formatting Tips section above. Break long sentences into shorter ones, add punctuation, and remove special characters. Regenerate with corrected text.
Issue: TTS mispronounces names or technical terms.
Cause: The TTS engine uses phonetic rules that may not match your intended pronunciation.
Solution: Try phonetic spelling (write the word as it sounds). For example, if the name “Siobhan” is mispronounced, try writing “Shivon” instead. Test multiple spelling variations until the output is correct.
Issue: Text is truncated in the generated audio.
Cause: The input text exceeds the 1,000-character limit.
Solution: Split the text into chunks of 1,000 characters or fewer. Generate each chunk separately, then combine the audio files. Ensure each chunk ends at a natural sentence boundary.
Next Steps: If the issue persists after these checks, visit the Support page with the error message and the text you attempted to generate.