Realtime transcription

Stream audio over a WebSocket and get transcripts back while the speaker is still talking.

The batch endpoint waits for a whole file. This one takes audio as it arrives and emits a transcript for each segment as soon as the speaker pauses — the same recogniser, a different rhythm.

Why this lives under /api, not /v1
The gateway won't carry this path: an Upgrade to /v1/realtime/transcribe is refused by LiteLLM, which owns /v1/realtime for its own OpenAI-realtime bridge and does not know this service's protocol. Verified both ways — the transcription service itself answers 101 Switching Protocols to the identical handshake. So live transcription is served same-origin, one level up, where the console proxies it through to the backend with the same key you already have.
The same WebSocket is served on both origins — wss://api.data.larsima.com/api/realtime/transcribe and wss://console.data.larsima.com/api/realtime/transcribe — with the identical contract and the same key. Use the api. host from server code; the console host is the same-origin one for pages served from it.
WS/api/realtime/transcribe
ParameterTypeDescription
auth
Optional
subprotocol | headerThe WebSocket subprotocol bearer.<key> is preferred — it's the only option a browser has (it cannot set a custom header on a WebSocket upgrade), and a subprotocol never lands in an access log or a Referer header the way a query parameter would. A server-side caller that isn't a browser can instead send a normal Authorization: Bearer <key> header, or X-Api-Key. There is no ?key=… query-string option.
binary frame
Optional
bytesRaw float32 mono PCM at 16kHz, or any container ffmpeg reads. Send audio in small chunks as it's captured, not the whole file at once — about 200 ms per frame is what the segmenter is tuned for (it counts each frame it receives as 0.2 s of audio when timing silence).
{"type":"config"}
Optional
text frameSets model and language before you send audio; answered with {"type":"ready",…}.
{"type":"commit"}
Optional
text frameFlush whatever is buffered right now, without waiting for silence.
{"type":"done"}
Optional
text frameEnd the session: flush anything left, then close after sending {"type":"done"} back.
A live session
# websocat is the practical "curl for WebSockets" -- there is no
# raw curl incantation that speaks this framing. Convert the file
# once, then stream the control frames and the raw audio together;
# -B keeps each write on its own frame instead of coalescing them.
ffmpeg -v error -i voice.m4a -ac 1 -ar 16000 -f f32le voice.raw

{ printf '%s\n' '{"type":"config","model":"whisper-large-v3","language":"persian"}'
  cat voice.raw
  printf '%s\n' '{"type":"done"}'
} | websocat -B 1000000 \
    --header "Sec-WebSocket-Protocol: bearer.$DATA_API_KEY" \
    wss://api.data.larsima.com/api/realtime/transcribe

# {"type":"ready","model":"whisper-large-v3","language":"persian"}
# {"type":"final","text":"قرارداد اجاره باید به صورت کتبی تنظیم شود.","duration":3.9}
# {"type":"done"}

Segmentation: cut on silence, not on a timer

A segment closes and a final event fires once at least a second of audio has arrived and the trailing 200ms has stayed below a silence threshold for 600ms straight, or once the buffer hits 20 seconds regardless of silence. The silence check looks only at the tail of the buffer, so a pause in the middle of a sentence doesn't end the segment early — only a pause that lasts.

What comes back

Only final events — there is no interim partial transcript while a segment is still filling. Each one carries text, duration (the length of that segment, in seconds — not its position in the overall stream), and language (the resolved ISO code for that segment). Like the batch endpoint, there are no word or segment timestamps beyond that.

A fourth message type, {"type": "error", "error": "…"}, can arrive instead — for an unknown model in config, or an unexpected failure partway through. The connection may still be open after one; treat it as informational, not necessarily as the end of the session. A missing or invalid key never gets this far: it is refused during the handshake itself (HTTP 403 on the upgrade, close code 1008), before any frame is exchanged. New connections are limited to 60 per minute per IP; there is no idle timeout on an open session.

model accepts the same three values as the batch endpoint — whisper-large-v3 (default), whisper-large-v3-turbo, whisper-fa. There is no fusion here; reconciling four transcripts costs roughly twice the time, which defeats the point of "while they're still talking".
Leave language out of config (or send "auto") and the first buffer long enough to trust is used to detect it — once that buffer reaches 3 seconds, the detected language is locked in for the rest of the session rather than re-detected on every segment; below 3 seconds, detection still runs but its answer is used for that one segment only. As with the batch endpoint, detection only overrides the Persian default when it's confident and not one of Persian's own confusables. Setting language explicitly in config skips detection entirely.
↑↓ Navigate ↵ Open esc Close