GEARBOXes FOR

AI inference

A model on a GPU is an engine and the inference stack around it is a gearbox. Using generic endpoints is like driving a car with its gearbox welded; you are either not driving the way you want or burning money running an inefficient engine. Zazu changes that. It offers stacks finetuned backwards from your objective to protect the metrics that matter while keeping the engine running optimally.

LLMs for voice agents are our first ready-to-use gearbox. It holds low latency throughout a call, paces its responses so your TTS can start sooner, and bills per minute, which undercuts any per-token bill. Model agnostic: tell us what you want or bring your own.

Live voice onboarding for production callers.

the coupling

Sheet 06: one GPU, vastly different operating modes. Illustrative.

More than an LLM API.

A voice agent has a shape and states: turns, interruptions, silences. All of these have a real impact on the engine's efficiency but an LLM API by itself sees none of that. Zazu enables easy coupling of the engine to the agent.

Sees
turns, interruptions, tool calls
Holds
GPU capacity, start of call to end
Adopts as
a Pipecat* plugin, listen-only, fail-safe
Speaks
an OpenAI-compatible LLM API

Simpler than it sounds: add the plugin, point it at the endpoint.

* Custom plugins available. LiveKit plugin coming soon.

What zazu's call-aware inference gets you

SHT 06 · NOTES

ITEM 01

Never fail mid-conversation

0 mid-call 429s

Zazu secures capacity when a call starts, or fails fast so you can react. A call in progress runs to the end.

ITEM 02

Starts warm, every call

no cold starts

Zazu recognises your system prompt and tool definitions. You are not billed tokens and time to keep reminding it.

ITEM 03

No awkward pauses

flat latency

Zazu holds latency steady from turn to turn. Barge-ins do not cost you billed but wasted tokens.

ITEM 04

Billed per minute, not per token

Zazu holds GPU capacity for each call so it can answer quickly and reliably. You pay for the minutes that capacity is held, not for the tokens that pass through it.

40%

cheaper than per-token billing, using tau-bench workload as a reference

Interested? Have a smoking engine that needs looking into?

000