VisemeTTS

Speech API for avatars

TTS that returns Oculus visemes with the audio.

One SSE stream: PCM speech, 15 Oculus visemes, and word clocks. Drive Streamoji, Ready Player Me, Unity, Unreal, or any rig that already speaks visemes.

sil
hello

Live viseme + word clocks, streamed with audio

Playground

Hear it. See the visemes. Read the word clocks.

Voices load from getVoices. Preview is a short clip. Speak streams avatar_ttsWithPoses so you get timings with the audio.

81/280

Ready.

sil
—
WordStartDur

Latest SSE audio event (chunk omitted) appears here after you speak.

How it works

Audio, visemes, and words in one stream.

01

Speak

POST text to the TTS endpoint with ttsEngineId visemetts and a voice id from getVoices.

02

Align

Word clocks come from the speech provider. MFA IPA (and rule-based G2P for OOV) maps phones onto those windows.

03

Stream

One SSE stream returns PCM audio, 15 Oculus visemes, and word timings. No separate aligner step.

Where it fits

Built for 3D avatars, games, and anything that already speaks visemes.

VisemeTTS is not a character platform. Bring your own avatar and LLM. We return the mouth shapes and clocks.

Streamoji & AiTwin avatars

Drop in ttsEngineId=visemetts on the existing avatar TTS path. 3D faces already consume Oculus visemes — no new rig.

Games & NPCs

Unity, Unreal, and Ready Player Me characters that speak Oculus visemes can lip-sync from the same SSE payload.

Virtual presenters

Word timings let you highlight captions, trigger gestures, and keep mouths locked to speech without a full NPC stack.

Alternatives

Why not Convai — and when we are better.

Convai is a strong NPC platform. VisemeTTS is the thinner layer: one stream you can plug into the avatar you already have. Same idea versus Azure or Polly viseme extras — streaming word clocks and an avatar-ready Oculus set in a single call.

VisemeTTSConvai
What you buyA speech API: audio + Oculus visemes + word timingsA full NPC stack: LLM, ASR, TTS, lip sync, knowledge, plugins
Lip sync15 Oculus visemes, streamed with the audioNeuroSync blendshapes or OVR visemes inside their character SDK
Your avatarStreamoji, Ready Player Me, custom Unity/Unreal rigsCharacters live on their platform
PricingPer character spokenPer interaction / quota
Best whenYou already have an avatar and only need aligned speechYou want a turnkey conversational NPC

API

One POST. SSE events. Visemes on every audio chunk.

Production calls use a client_* developer JWT. The playground on this site is origin-limited and unauthenticated.

Full contract
curl -N https://ai.aitwin.me/avatar_ttsWithPoses \
  -H 'Content-Type: application/json' \
  -H 'Authorization: Bearer client_<jwt>' \
  -d '{
    "user_query": "Hello from VisemeTTS.",
    "ttsEngineId": "visemetts",
    "voice_id": "eve"
  }'

FAQ

What viseme set do you return?+

The 15 Oculus visemes: sil, PP, FF, TH, DD, kk, CH, SS, nn, RR, aa, E, I, O, U. Each has start and duration in seconds, chunk-relative on audio events.

Do I get word timings too?+

Yes. Each audio event includes words with start, duration, phones, and an OOV flag when G2P was used.

Is this a Convai alternative?+

For lip-synced speech, yes. Convai is a full conversational NPC platform. VisemeTTS is only TTS + visemes + words, so you keep your avatar, game engine, and LLM.

How is usage billed?+

Per character spoken, same ladder as AiTwin: 10k free, then 200k for $19/month, and up. Unused credits do not roll over.