File transcription
One JSON result
Upload a recording and receive the complete transcript when processing finishes.
Speech-to-Text model
Listening is not waiting for silence. It is knowing what is tentative, what has been committed and when the next part of a conversation is ready to begin.
The gateway checks these values before a single second of audio is decoded, so an integration can validate against them on the client and fail early.
Loli 2.0 accepts both complete recordings and continuous microphone audio. The model stays the same while the transport changes how soon your application can observe the result.
One JSON result
Upload a recording and receive the complete transcript when processing finishes.
NDJSON by chunk
Follow a long recording line by line, including started, chunk, final and error objects.
WebSocket events
Send PCM16 continuously and receive replaceable partials, committed segments and a final transcript.
Automatic mode uses server voice activity detection to recognise speech and silence. Manual mode waits for your application to commit, which fits Push-to-Talk and clients that already own their turn detection.
Stream audio continuously. Partials evolve while the person speaks, and segments commit at detected silence boundaries.
Buffer audio until the client sends commit. Each valid commit produces one final segment for that turn.
Important
Use auto when the model should detect among supported languages. Choose one of the published single-language modes when context is known, or a bilingual mode when a conversation naturally crosses a supported pair.
A partial is a live hypothesis and may be replaced. A segment commits one utterance. Final represents the complete session transcript. Treating those events differently keeps live captions responsive without mistaking provisional words for settled input.
partialReplaceable text for the utterance currently in progress.
segmentA committed utterance produced at a boundary or manual commit.
finalThe complete transcript emitted when the session stops.
closedConfirmation that the session has finished and released its state.
File transcription recognises the published audio containers after decoding. Real-time sessions receive signed 16-bit little-endian mono PCM, while the start message declares the client sample rate so the service can normalise it.
Transcribe a familiar clip first. When the words look right, move the same product idea into NDJSON progress or a real-time session and design how provisional text becomes action.
Was this page helpful?