The sessions page in Anam Lab shows a
Performance card with four latency metrics for your recent sessions. Each
metric shows the median (p50), p95, and maximum across every conversational
turn in the selected time period, with a green, amber, or red indicator next
to it.
Latency data comes from session reports, which are produced for sessions
where Anam runs the conversational pipeline (turnkey, custom LLM, and
ElevenLabs agent sessions). You can fetch the same data programmatically with
the aggregate analytics
endpoint or drill
into a single session with the per-session analytics
endpoint.
Audio-passthrough and LiveKit sessions do not produce these conversational reports. A session count can include sessions without reports; missing stage timings are not zero latency. For an external LLM/TTS pipeline, instrument your own provider and client path instead.
What the indicators mean
The indicator compares your median (p50) for the selected period against the following thresholds.
These indicators are diagnostic thresholds, not end-to-end latency guarantees for a particular country, browser, or integration.
- Green — at or better than a healthy production median. No action
needed.
- Amber — slower than typical. Users will notice the delay; review the
suggestions below.
- Red — slower than roughly 90% of production sessions. Something in
your configuration or environment is adding significant delay.
Response latency
Response latency measures the reported interval from user speech end to persona speech start. These are pipeline timestamps, not a measurement of when sound reaches the user’s speakers or a speech frame reaches their screen. Start by checking which stage is amber or red, but measure client playback separately when evaluating the user’s experience.
If all three stages are green but response latency is amber or red, the
gap needs investigation rather than being attributed to one cause:
- Check the actual served region. Automatic routing uses location and routing policy; it does not guarantee the lowest-latency route.
- Use the per-session analytics
endpoint to find the slow
sessions, then check their
clientMetadata and location for a pattern.
Transcription
Transcription latency is how long it takes to produce the final transcript
after the user stops speaking. Anam manages the transcription pipeline, so
this stage is normally well under half a second.
If it is amber or red:
- Review your voice detection settings.
A high
endOfSpeechSensitivity can delay the point at which Anam decides
the user has finished speaking.
- Check whether the affected sessions come from users on poor network
connections; delayed or dropped audio extends transcription time.
- If it stays red across many sessions with default settings, contact your
Anam representative or support — this stage is Anam-managed.
LLM first output
LLM first output is the time from transcript completion to the first token
from the language model. This is usually the largest and most controllable
stage.
If it is amber or red:
- Choose a faster model. Larger models take longer to produce their
first token. See available LLMs for the
options and their trade-offs.
- Shorten your system prompt. Very long prompts increase time to first
token. The prompting guide covers how to
keep prompts effective and compact.
- If you use a custom LLM, host it close to Anam’s infrastructure,
make sure streaming is enabled, and measure your endpoint’s own time to
first token — Anam can only be as fast as your server. See
custom LLMs.
- Review knowledge and tools. Knowledge retrieval and tool calls run
during the response and add to the delay on turns that use them. Check
the tool call timings on the
session endpoint to see how long
each call took.
TTS first audio
The Lab metric called TTS first audio uses the reported interval from TTS start to persona speech start. It is not the same as your TTS provider’s first response byte, nor does it measure browser playback. Compare measurements only when their start and end events match.
If it is amber or red:
- Try a different voice. Some voices and voice providers generate the
first audio chunk faster than others; see
voice configuration.
- If it stays red across voices, contact your Anam representative or
support — this stage is Anam-managed.
Measure your own integration
Keep these measurements separate:
- Provider first byte: the provider request to its first response byte; that byte may not yet contain decoded PCM.
- First audio submitted: when your application first calls
sendAudioChunk() with usable PCM. This does not acknowledge delivery to the engine.
- Initial displayed video frame: the first video-frame callback in the client. It can be an idle frame, not the first frame of a particular utterance.
- Conversational latency: the interval from the user’s speech end to the matching response being heard or displayed on the client. Provider timing and engine timing alone do not measure this interval.
In JavaScript SDK 4.27.0, register a VIDEO_PLAY_STARTED listener before starting the stream to observe its initial video-frame milestone. The SDK uses HTMLVideoElement.requestVideoFrameCallback() for this event. It is not a per-utterance speech-start or playback-finished event; do not subtract every audio submission timestamp from it to claim a conversational latency result.
For a reproducible evaluation, record the SDK/model configuration, served region, provider, browser/device, network, and exact timing boundaries. Use a monotonic clock such as performance.now() for elapsed times within one client; do not subtract unsynchronized browser and server clocks. Repeat measurements and report their distribution rather than a single best result. A per-utterance audio-submission-to-speech-frame measurement needs a validated way to match the displayed response to that utterance; the initial-frame event alone does not provide it.
Browser and mobile checks
WebRTC capability alone does not establish that every browser version is supported or tested. Check the APIs used by your integration, including requestVideoFrameCallback() for the SDK’s video-element playback path. Microphone capture requires a secure context and permission. Test audio playback from a user gesture, iframe microphone permissions, device changes, and background/resume behavior on the actual mobile Safari or Chrome versions you deploy.
For the underlying browser API requirements, see video-frame callbacks and microphone capture. These API references are not an Anam-tested browser/version matrix. Confirm any required support matrix or regional latency commitment with Anam before committing it to your users.
- Aggregate analytics
endpoint —
the same percentiles as the Performance card, filterable by persona, API
key, client label, and session type, including the slowest individual
turns in the period.
- Per-session analytics
endpoint — per-turn
timing for one session, so you can see exactly which stage was slow on
which turn.
Last modified on September 16, 2026