A Retell AI web voice agent can share an interface with a real-time avatar, but the current public Web SDK does not document a production-ready stream of contiguous assistant PCM for motion generation. A reliable Spatius deployment therefore needs a supported assistant-audio handoff—or application-owned final TTS audio—before it should be called an integration.
Independent guide: Retell AI and Spatius are separate products. Spatius is not affiliated with, endorsed by, or an official integration partner of Retell AI. This article does not claim a native connector.
Last verified: September 24, 2026.
Retell handles the voice agent; Spatius handles the visual layer
Retell’s web deployment documentation describes browser voice calls over WebRTC, alongside text chat and phone callbacks. Retell owns the conversation, speech pipeline, tools, call state, and audio playback. Spatius can add a locally rendered character when the application can provide the assistant speech that the user actually hears.
That boundary matters. A phone-only agent has no visual surface, so it normally gains nothing from an avatar. A Retell agent embedded in a product, kiosk, training interface, or guided workflow can benefit from a visible character—but talking state alone is not enough for speech-driven facial motion.
| Question | Current public evidence | Production implication |
|---|---|---|
| Can Retell run a web voice call? | Yes | A visible web experience is possible |
| Does the SDK expose talking state? | Yes | Useful for UI state, not detailed motion |
| Can it emit audio analyser samples? | Yes, when enabled | Useful for meters and reactive visualizers |
| Are those samples contiguous PCM frames? | No; the SDK source identifies them as snapshots | Do not treat them as the complete assistant utterance |
| Is a native Spatius connector documented? | No | Treat the architecture as conditional until an approved audio path exists |
What the Retell Web SDK actually exposes
Retell’s audio basics documentation explains that the frontend Web SDK abstracts browser audio handling. The open-source RetellWebClient can enable emitRawAudioSamples, create an analyser for the remote agent track, and emit Float32Array snapshots during animation frames. The source explicitly notes that these are not contiguous PCM frames.
That makes the event appropriate for a waveform, volume-reactive glow, or agent-speaking indicator. It is not sufficient evidence for accurate speech-driven avatar motion. Analyser windows may overlap, skip time, or represent sampled visualization data rather than the complete ordered speech played to the user.
Spatius expects the approved assistant speech in order. Its audio guide identifies mono signed 16-bit little-endian PCM as the canonical audio representation and distinguishes ordinary end-of-input from interruption.
Where Spatius would fit
If a supported assistant-audio boundary is available, the architecture remains modular:
User microphone
↓
Retell web voice session
├─ conversation, tools, call state, and interruptions
└─ supported final assistant-audio handoff
↓
application coordinator
├─ authoritative playback path
└─ ordered avatar-audio path
↓
Spatius Motion Server
↓ motion data
AvatarKit client render
Retell would continue to own the voice agent. Spatius would receive only the assistant speech required to generate motion and would render the character locally. Prompts, user audio, private transcripts, tool arguments, and internal call state would not need to be sent to Spatius merely to animate the avatar.
Three possible architecture paths
1. A supported contiguous assistant-audio interface
This is the preferred path. Retell or another approved interface provides the actual ordered assistant audio before or during playback. The application converts it once, keeps one authoritative playback path, and sends the matching utterance to Spatius Motion Server.
Before implementation, ask Retell to confirm the supported interface, codec, sample rate, channel count, frame ordering, response boundaries, interruption semantics, and whether the audio may be routed to a separate renderer.
2. Application-owned final TTS audio
If the application legitimately owns the final voice bytes before playback, it may be able to fan out identical speech to the user and Spatius. This path depends on the actual Retell architecture and commercial terms. It must not be assumed from the standard Web SDK, and it should not be created by scraping private SDK internals.
3. Analyser-driven reactive animation
The public analyser snapshots can support scale, glow, idle-versus-speaking state, or a waveform. That can be a useful product treatment, but it should be labeled as reactive visualization rather than accurate speech-driven facial motion.
| Path | Motion fidelity | Integration confidence | Recommendation |
|---|---|---|---|
| Supported contiguous assistant audio | Can support speech-driven motion | High after format and lifecycle validation | Preferred |
| Application-owned final TTS audio | Can support speech-driven motion | Depends on real architecture and permission | Evaluate carefully |
| SDK analyser snapshots | Coarse reactive state only | Documented for visualization | Do not present as full lip sync |
A safe proof-of-concept plan
- Build the Retell web call and validate ordinary playback, completion, and interruption.
- Obtain written confirmation of a supported contiguous assistant-audio source.
- Keep long-lived Retell and Spatius credentials on the server; issue only short-lived client credentials.
- Assign one turn ID to the Retell response, playback queue, and Spatius input.
- Convert the supported speech once and preserve chunk order and response boundaries.
- On interruption, stop playback and clear the matching Spatius turn together.
- Log first audio, first motion, dropped or late chunks, reconnects, and cleanup.
Do not begin by attaching motion generation to the analyser event. First prove that the application has the complete speech source and the right to use it for this purpose.
What Spatius contributes after the audio boundary is solved
Once the application can supply approved assistant speech, Spatius Motion Server can return motion data while AvatarKit renders the avatar locally. This is different from sending a second cloud-rendered video stream through the Retell call.
Local rendering gives the product control over avatar placement, camera, captions, accessibility controls, tool-result UI, loading states, and audio-only fallback. It also keeps measurement boundaries visible:
| Measurement | Retell/application layer | Spatius/avatar layer |
|---|---|---|
| Call setup and time to first assistant audio | Measure here | Not included |
| Agent reasoning, tools, and voice playback | Measure here | Not included |
| Audio handoff completeness | Application boundary | Required input validation |
| Audio-to-motion delivery | Not included | Measure here |
| Client rendering stability | Product client | Measure AvatarKit on target devices |
| Interruption correctness | Shared product event | Must stop with playback |
Read the broader real-time avatar integration patterns before choosing a transport and recovery model.
When this combination is a good fit
Consider Retell plus Spatius when all of the following are true:
- the experience has a screen and a visible character improves the workflow;
- Retell remains the desired voice-agent platform;
- a supported complete assistant-audio path is available;
- the product wants local avatar rendering rather than a finished remote video stream;
- the team can coordinate interruption and recovery across two services.
Keep the experience voice-only when it is primarily a phone call, when the analyser snapshots are the only accessible audio signal, or when the team wants one vendor to own the complete agent-and-avatar stack.
Cost and buying decision
Retell publishes usage-based voice-agent pricing and component charges on its official pricing page. Spatius and application infrastructure are separate layers, so compare them independently using Spatius pricing. Do not publish one blended rate without documenting the selected Retell components, avatar usage, concurrency, and infrastructure assumptions.
For a worked pattern where the platform documents explicit output-audio events, compare the OpenAI Realtime API avatar tutorial.
Retell AI avatar FAQ
Is there an official Retell AI and Spatius integration?
This guide does not claim one. It describes the requirements for a third-party avatar handoff and explains why a supported complete assistant-audio stream is necessary.
Can Retell's audio event be sent directly to Spatius?
Not as a production assumption. The public SDK source describes analyser snapshots rather than contiguous PCM frames for the complete assistant utterance.
Can talking-state events drive lip sync?
They can switch idle and speaking states, but they do not contain enough speech information for accurate speech-driven facial motion.
Does Spatius replace Retell AI?
No. Retell remains the voice-agent platform. Spatius would use approved assistant speech to generate motion and render the avatar locally.
What should a team ask Retell before building?
Ask for a supported assistant-audio interface, format and ordering guarantees, response and interruption semantics, and permission to route that audio to a separate renderer.
Validate the audio boundary before promising an avatar
The honest decision is conditional: Retell can remain the agent and Spatius can become the visual layer, but only after the application has a supported complete assistant-audio path. That boundary should be proven before the experience is described as integrated.
Bring your Retell web architecture and available audio interfaces. We will help you determine whether the current boundary can support synchronized Spatius motion without undocumented SDK internals. Request a demo, or ,或Review Spatius pricing.。