Never fail mid-conversation
0 mid-call 429sZazu secures capacity when a call starts, or fails fast so you can react. A call in progress runs to the end.
A model on a GPU is an engine and the inference stack around it is a gearbox. Using generic endpoints is like driving a car with its gearbox welded; you are either not driving the way you want or burning money running an inefficient engine. Zazu changes that. It offers stacks finetuned backwards from your objective to protect the metrics that matter while keeping the engine running optimally.
LLMs for voice agents are our first ready-to-use gearbox. It holds low latency throughout a call, paces its responses so your TTS can start sooner, and bills per minute, which undercuts any per-token bill. Model agnostic: tell us what you want or bring your own.
Live voice onboarding for production callers.
the coupling
Sheet 06: one GPU, vastly different operating modes. Illustrative.
A voice agent has a shape and states: turns, interruptions, silences. All of these have a real impact on the engine's efficiency but an LLM API by itself sees none of that. Zazu enables easy coupling of the engine to the agent.
Simpler than it sounds: add the plugin, point it at the endpoint.
* Custom plugins available. LiveKit plugin coming soon.
SHT 06 · NOTES
Zazu secures capacity when a call starts, or fails fast so you can react. A call in progress runs to the end.
Zazu recognises your system prompt and tool definitions. You are not billed tokens and time to keep reminding it.
Zazu holds latency steady from turn to turn. Barge-ins do not cost you billed but wasted tokens.
Zazu holds GPU capacity for each call so it can answer quickly and reliably. You pay for the minutes that capacity is held, not for the tokens that pass through it.
40%
cheaper than per-token billing, using tau-bench workload as a reference
Interested? Have a smoking engine that needs looking into?
000