Voice overview
Connect a voice agent or the microphone, and the visual follows the conversation — its state and its sound.
A voice source turns a conversation into two streams the visual understands:
- the agent's state (
initializing,idle,listening,thinking,speaking), which picks the look through your states; - the audio: an overall level and per-band levels, which drive
audioLevel(an orb's breathing) and the bars and waveforms of a signal.
Give the view a source and it does the rest: it follows the state, eases the audio at display rate so nothing steps, shows the mute cue when you set it, and flashes briefly when the user interrupts the agent (barge-in).
The two talking states move in opposite directions, which is what makes them tell apart at a glance: while the user speaks, the visual is drawn inward by the mic, and while the agent speaks it swells outward with its voice. That behaviour ships with the engine, so a visual follows a conversation without a design file; an FX Spec then changes whatever you want on top of it.
import { mount } from "@sinua/web";
import { LocalMicVoiceSource } from "@sinua/voice/mic";
import voiceOrb from "../spec/voice-orb.fxspec.json";
const canvas = document.querySelector<HTMLCanvasElement>("#orb")!;
const voice = new LocalMicVoiceSource();
// The view binds the source: level and bands drive the visual, and the source's
// AgentState picks the spec's lifecycle state. It never connects the source itself.
const fx = mount(canvas, { spec: voiceOrb, voice });
// Connect from a user gesture: the mic prompt and audio playback both need one.
document.querySelector("#talk")!.addEventListener("click", async () => {
try {
await voice.connect();
} catch (err) {
console.error("voice failed to start", err); // e.g. the mic permission was denied
}
});
// On teardown: voice.disconnect(); fx.destroy();
export { fx, voice };
Sources
| Source | Connects to | Page |
|---|---|---|
| LiveKit | an existing LiveKit room with an agent | LiveKit |
| OpenAI Realtime | OpenAI's Realtime API over WebRTC | OpenAI Realtime |
| Gemini Live | Google's Gemini Live API | Gemini Live |
| ElevenLabs | ElevenLabs Conversational AI | ElevenLabs |
| Microphone, test tone | the device mic, or a synthetic voice for previews | Mic & test tone |
Each vendor source is its own package on iOS and Android, and its own entry point of @sinua/voice on the web (/livekit, /openai, /gemini, /elevenlabs, /mic, /tone), so your app only includes the SDKs it uses. On the web, install @sinua/voice (and livekit-client only if you use LiveKit), pass the source to the view as voice, and call connect() from a user gesture, such as a button tap: browsers only allow audio and the microphone after one. All of them report the same state names, the same audio shape, and interruption where the vendor signals it.
Credentials: never ship an API key
A voice source never needs your vendor API key. Keep the key on your server, mint a short-lived credential per session there, and hand only that credential to the source:
- OpenAI: an ephemeral client secret (
ek_…), minted withPOST /v1/realtime/client_secrets. - Gemini: an ephemeral auth token (
auth_tokens/…). - ElevenLabs: a signed URL for a private agent, or just the agent id for a public one.
- LiveKit: an access token from your backend, or a room you've already connected.
Sinua doesn't ship a backend. The mint endpoint is yours, so you decide who gets a session and with which model, voice and instructions.
Driving the visual yourself
If your app already listens to the agent, skip the source: pass a VoiceOverrides you feed yourself as voice. Call push({ level, bands }) for the audio, setState(...) for the agent's state and interrupt() for a barge-in. The view's easing and cues still apply.
The mute cue is yours to turn on, since sources don't report mute: voiceOptions: { muted: true } on the web, or the muted option of VoiceOverrides on iOS and Android.