Available Models & Agents
Browse 29 models and agents with text-to-speech capabilities
mayaresearch/veena-max
mayaresearch
VeenaMAX is a high-performance, advanced Text-to-Speech API that converts text into natural-sounding speech with emotional intelligence and blazing-fast response times. It can speak English and Hindi - the most widely used languages in India. As a signifi
inworld/realtime-tts-2
inworld
Most expressive text-to-speech model from Inworld, with natural-language steering, real-time latency, and multilingual support across 100+ languages.
inworld/realtime-tts-1.5-max
inworld
Highest-quality realtime text-to-speech with <200ms latency, emotion control, and 15-language support
inworld/realtime-tts-1.5-mini
inworld
Ultra-fast, cost-efficient realtime text-to-speech with ~120ms latency and 15-language support
xai/grok-text-to-speech
xai
Convert text to natural-sounding speech with xAI's Grok TTS. 5 voices, 20 languages, expressive speech tags, and high-fidelity MP3 / WAV / telephony audio output.
google/gemini-3.1-flash-tts
Google's fast, expressive text-to-speech model with 30 voices and 70+ language support
adirik/hierspeechpp
adirik
Zero-shot speech synthesizer for text-to-speech and voice conversion
lucataco/whisperspeech-small
lucataco
An Open Source text-to-speech system built by inverting Whisper
cjwbw/melotts
cjwbw
High-quality multilingual text-to-speech library
cjwbw/parler-tts
cjwbw
lightweight text-to-speech (TTS) model, trained on 10.5K hours of audio data
zsxkib/hololive-style-bert-vits2
zsxkib
🎙️Hololive text-to-speech and voice-to-voice (Japanese🇯🇵 + English🇬🇧)
lee101/guided-text-to-speech
lee101
Guided Text to Speech Generator
e1100x/chattts
e1100x
ChatTTS is a text-to-speech model designed specifically for dialogue scenarios such as LLM assistant.
jaaari/kokoro-82m
jaaari
Kokoro v1.0 - text-to-speech (82M params, based on StyleTTS2)
alphanumericuser/kokoro-82m
alphanumericuser
Kokoro v1.0 - text-to-speech (82M params, based on StyleTTS2)
cuuupid/zonos
cuuupid
Zonos-v0.1 beta, a SOTA text-to-speech Transformer model with extraordinary expressive range, built by Zyphra.
lucataco/step-audio-tts-3b
lucataco
Step-Audio-TTS-3B represents the industry's first Text-to-Speech (TTS) model trained on a large-scale synthetic dataset utilizing the LLM-Chat paradigm
cjwbw/voicecraft
cjwbw
Zero-Shot Speech Editing and Text-to-Speech in the Wild
lucataco/higgs-audio-v2
lucataco
Higgs Audio v2, a powerful text-to-speech audio foundation model that excels in expressive audio generation
microsoft/vibevoice
microsoft
Microsoft's VibeVoice text-to-speech model that can generate long-form speech from text with sample voices.