Speak
POST text to the TTS endpoint with ttsEngineId visemetts and a voice id from getVoices.
Speech API for avatars
One SSE stream: PCM speech, 15 Oculus visemes, and word clocks. Drive Streamoji, Ready Player Me, Unity, Unreal, or any rig that already speaks visemes.
Live viseme + word clocks, streamed with audio
Playground
Voices load from getVoices. Preview is a short clip. Speak streams avatar_ttsWithPoses so you get timings with the audio.
Ready.
| Word | Start | Dur |
|---|
Latest SSE audio event (chunk omitted) appears here after you speak.
How it works
POST text to the TTS endpoint with ttsEngineId visemetts and a voice id from getVoices.
Word clocks come from the speech provider. MFA IPA (and rule-based G2P for OOV) maps phones onto those windows.
One SSE stream returns PCM audio, 15 Oculus visemes, and word timings. No separate aligner step.
Where it fits
VisemeTTS is not a character platform. Bring your own avatar and LLM. We return the mouth shapes and clocks.
Drop in ttsEngineId=visemetts on the existing avatar TTS path. 3D faces already consume Oculus visemes — no new rig.
Unity, Unreal, and Ready Player Me characters that speak Oculus visemes can lip-sync from the same SSE payload.
Word timings let you highlight captions, trigger gestures, and keep mouths locked to speech without a full NPC stack.
Alternatives
Convai is a strong NPC platform. VisemeTTS is the thinner layer: one stream you can plug into the avatar you already have. Same idea versus Azure or Polly viseme extras — streaming word clocks and an avatar-ready Oculus set in a single call.
| VisemeTTS | Convai | |
|---|---|---|
| What you buy | A speech API: audio + Oculus visemes + word timings | A full NPC stack: LLM, ASR, TTS, lip sync, knowledge, plugins |
| Lip sync | 15 Oculus visemes, streamed with the audio | NeuroSync blendshapes or OVR visemes inside their character SDK |
| Your avatar | Streamoji, Ready Player Me, custom Unity/Unreal rigs | Characters live on their platform |
| Pricing | Per character spoken | Per interaction / quota |
| Best when | You already have an avatar and only need aligned speech | You want a turnkey conversational NPC |
API
Production calls use a client_* developer JWT. The playground on this site is origin-limited and unauthenticated.
Full contractcurl -N https://ai.aitwin.me/avatar_ttsWithPoses \
-H 'Content-Type: application/json' \
-H 'Authorization: Bearer client_<jwt>' \
-d '{
"user_query": "Hello from VisemeTTS.",
"ttsEngineId": "visemetts",
"voice_id": "eve"
}'The 15 Oculus visemes: sil, PP, FF, TH, DD, kk, CH, SS, nn, RR, aa, E, I, O, U. Each has start and duration in seconds, chunk-relative on audio events.
Yes. Each audio event includes words with start, duration, phones, and an OOV flag when G2P was used.
For lip-synced speech, yes. Convai is a full conversational NPC platform. VisemeTTS is only TTS + visemes + words, so you keep your avatar, game engine, and LLM.
Per character spoken, same ladder as AiTwin: 10k free, then 200k for $19/month, and up. Unused credits do not roll over.