SUM AI Voice, a real-time voice assistant on the website

The problem

We already run an AI text chat on this site. It works, but typing is friction. A visitor with a real question has to phrase it, type it, wait, read, and type again. Talking is faster. The question was whether a website assistant could actually hold a spoken conversation that felt live, rather than the stilted press-to-talk, wait-for-the-beep experience most voice bots ship.

"Real time" is the hard part. A natural conversation has sub-second turn-taking, and people interrupt each other constantly. The moment you can hear an answer forming and want to cut in with "no, the other one," a voice assistant that can't be interrupted feels broken. So the bar we set was simple to state and hard to hit: continuous listening, spoken replies within a fraction of a second, and the ability to talk over the assistant at any time.

What SUM AI Voice does

SUM AI Voice is a spoken front door to the company. A visitor opens it, presses Start, and asks whatever they want to know about SUM AI: what we build, how we work, whether we can help with their problem. It answers out loud, in real time.

Key capabilities:

  • Continuous, hands-free listening. No push-to-talk. The mic streams the whole time; the model decides when you have finished a thought and responds.
  • Barge-in. You can talk over the assistant and it stops immediately, the way a person would. No waiting for it to finish a sentence you have already answered.
  • Native spoken answers. Replies are generated as audio directly, not typed out and read by a separate text-to-speech voice. The result sounds noticeably more natural, with real intonation and pacing.
  • Live transcripts. Both sides of the conversation are transcribed and shown on screen as they happen, so the exchange is readable as well as audible.
  • Grounded answers. It speaks from a curated description of what SUM AI actually does, and is told to be honest about what it does not know rather than invent specifics.
  • Understands any language, answers in English. It can follow a visitor speaking another language, but keeps its spoken replies in clear English.

Architecture

The whole thing is three moving parts: the browser, a small server we run, and the Gemini Live API. The browser captures the microphone and plays audio back. The Live API does the listening, thinking, and speaking. The server in the middle exists for one specific reason.

A browser cannot open an authenticated WebSocket to Google directly, because the WebSocket API in the browser has no way to set an Authorization header, and we are never going to ship a cloud credential into client-side JavaScript anyway. So a thin proxy sits in between: the browser connects to our server over a WebSocket, our server holds the credential, opens the upstream connection to the Live API, sends the session's setup message, and then relays audio frames in both directions.

Stage What happens
Connect Browser opens a WebSocket to our proxy. The proxy authenticates to Google, opens the upstream Live socket, and sends the setup message: model, voice, response format, and the system prompt.
Talk Mic audio is downsampled to 16 kHz mono PCM in the browser and streamed up as a continuous series of small chunks.
Listen The model streams 24 kHz PCM audio back, chunk by chunk. The browser schedules the chunks back-to-back so playback is gapless.
Interrupt If the visitor starts speaking again, an interruption signal tells the browser to drop whatever audio is still queued so the assistant goes quiet at once.
Transcript Text transcriptions of both sides arrive alongside the audio and are rendered live.
Browser ↔ proxy ↔ Gemini Live API. The proxy holds the credential and relays frames verbatim in both directions.

Keeping the proxy "dumb" is deliberate. It does not transcode audio, rewrite messages, or buffer turns. It injects the setup once and then passes frames through. That keeps latency low and makes the data path easy to reason about: the model is doing the real work, and the proxy is just the part of the pipe that can safely hold a credential.

Native audio, not text-to-speech

The model is gemini-live-2.5-flash-native-audio, with the Orus voice. The word that matters there is native. A common way to build a voice bot is to run a normal text model and pipe its words through a separate text-to-speech engine. It works, but it sounds like it works that way: flat, evenly paced, a beat behind. A native-audio model generates the speech directly, so intonation, emphasis, and rhythm come from the model itself rather than being bolted on afterward.

Practically, it also collapses two network round-trips into one. There is no "generate text, then send text to a voice service, then stream audio" chain. The model hears audio and emits audio, which is a big part of why the replies land fast enough to feel conversational.

The audio pipeline

Audio in and audio out run at different sample rates, and both happen in the browser with the Web Audio API and no build step. Going up, the microphone is captured, mixed to mono, and downsampled to 16 kHz, the rate the model expects for input. Each chunk is converted to 16-bit PCM, base64-encoded, and streamed as a real-time audio frame. We also turn on the browser's echo cancellation, noise suppression, and auto gain on the mic so the assistant does not hear itself and interrupt its own answer.

Coming back down, the model streams 24 kHz PCM. The naive approach, playing each chunk the moment it arrives, produces clicks and gaps, because chunks do not arrive on a perfectly even cadence. Instead we keep a running "play head" timestamp and schedule each incoming chunk to start exactly where the previous one ends. The audio plays seamlessly even though it is arriving in pieces over the network.

Barge-in: talking over the assistant

This is the feature that separates a conversation from a kiosk. The Live API runs voice-activity detection on the incoming stream, so it knows when the visitor has started speaking again, even while the assistant is mid-reply. When that happens, it sends an interruption signal.

The interruption itself is simple to handle but easy to get wrong. Because we scheduled audio chunks ahead of time for gapless playback, there can be several seconds of speech already queued to play. On an interruption we stop every queued audio source immediately and reset the play head, so the assistant goes quiet the instant you start talking, not after it finishes the sentence it had buffered. Without that, "interrupting" would feel laggy and the whole illusion would break.

Live transcripts

Voice alone is fragile. You mishear a name, you want to re-read a detail, or you are somewhere you can't use sound. So the session asks the Live API to transcribe both sides, and the proxy relays those transcripts to the browser as they stream. Each turn is rendered as a chat-style bubble that fills in as the words arrive. You get the immediacy of voice with the scannability of text, and the on-screen record makes the whole thing feel grounded rather than ephemeral.

Giving it a personality that stays honest

A voice assistant's "prompt" is its personality, and speaking out loud changes the rules. Long, formatted, list-heavy answers that read fine on a page are exhausting to listen to. So the assistant is told to speak in short, conversational turns, usually a sentence or two of plain spoken prose, no markdown, no bullet points, web addresses and numbers said the way a person would say them.

More importantly, it is grounded. It speaks from a curated description of what SUM AI actually does, and it is explicitly told to admit when it does not know something rather than invent a price, a timeline, or a guarantee. For a voice on the company's own website, sounding confident matters far less than never saying something untrue, so the prompt optimizes for honesty first and polish second.

Engineering decisions

  1. Native audio beats text-to-speech for naturalness. Letting one model both understand and speak removes a whole TTS hop and produces intonation that a text-then-speech pipeline can't match.
  2. A thin proxy is the right place for the credential. Browsers can't authenticate a WebSocket to the model and shouldn't hold cloud keys. A pass-through proxy that injects setup once keeps secrets server-side without adding latency.
  3. Schedule playback, don't just play it. Tracking a play head and scheduling chunks back-to-back is the difference between gapless speech and a stream of clicks.
  4. Barge-in means flushing the queue, not just muting. Because audio is buffered ahead, an interruption has to actively stop the queued sources, or the "interrupt" feels seconds late.
  5. Echo cancellation is not optional. With the speaker and mic both live, the assistant will interrupt itself unless the browser's echo cancellation and noise suppression are on. Headphones make it flawless.
  6. Write the prompt for the ear, not the eye. Spoken replies need to be short and free of formatting, and the assistant has to be told to stay honest, admitting the unknown instead of inventing specifics.

Try it

The best way to understand a voice assistant is to talk to it. Open SUM AI Voice, press Start, and ask it what we do, out loud. It will answer the same way. If you would rather talk to a human, that option is one message away.