Speech-to-Text model

Loli 2.0: Turn sound into something your product can understand

Listening is not waiting for silence. It is knowing what is tentative, what has been committed and when the next part of a conversation is ready to begin.

On this page

The contract in one screen

The gateway checks these values before a single second of audio is decoded, so an integration can validate against them on the client and fail early.

Modality
Speech to text
Transports
REST, NDJSON, WebSocket
File formats
WAV, MP3, WEBM, OGG, M4A, MP4
Upload ceiling
100 MB
Audio duration
0.2 s - 1,800 s
Real-time input
PCM16 mono, 8,000 - 192,000 Hz
Languages
30 plus auto, 3 bilingual
Session events
partial, segment, final, closed

Choose how the model listens

Loli 2.0 accepts both complete recordings and continuous microphone audio. The model stays the same while the transport changes how soon your application can observe the result.

File transcription

One JSON result

Upload a recording and receive the complete transcript when processing finishes.

Progressive transcription

NDJSON by chunk

Follow a long recording line by line, including started, chunk, final and error objects.

Real-time listening

WebSocket events

Send PCM16 continuously and receive replaceable partials, committed segments and a final transcript.

Let the turn boundary match the product

Automatic mode uses server voice activity detection to recognise speech and silence. Manual mode waits for your application to commit, which fits Push-to-Talk and clients that already own their turn detection.

Automatic

Stream audio continuously. Partials evolve while the person speaks, and segments commit at detected silence boundaries.

Manual

Buffer audio until the client sends commit. Each valid commit produces one final segment for that turn.

Important

Send stop before closing so the server can emit the final transcript followed by the closed event.

Listen broadly, or listen with a deliberate constraint

Use auto when the model should detect among supported languages. Choose one of the published single-language modes when context is known, or a bilingual mode when a conversation naturally crosses a supported pair.

  • Automatic language detection
  • 30 published single-language modes
  • English and Vietnamese
  • Korean and English
  • Japanese and English

Build the interface around event meaning

A partial is a live hypothesis and may be replaced. A segment commits one utterance. Final represents the complete session transcript. Treating those events differently keeps live captions responsive without mistaking provisional words for settled input.

partial

Replaceable text for the utterance currently in progress.

segment

A committed utterance produced at a boundary or manual commit.

final

The complete transcript emitted when the session stops.

closed

Confirmation that the session has finished and released its state.

Give the model clean, explicit audio

File transcription recognises the published audio containers after decoding. Real-time sessions receive signed 16-bit little-endian mono PCM, while the start message declares the client sample rate so the service can normalise it.

  • WAV 16-bit PCM, MP3, WebM, OGG, MP4 or M4A for files
  • PCM16 little-endian mono for real-time input
  • Declared input sample rate from 8,000 to 192,000 Hz
  • Send stop before closing a real-time session

Start with one recording you know by heart

Transcribe a familiar clip first. When the words look right, move the same product idea into NDJSON progress or a real-time session and design how provisional text becomes action.

Was this page helpful?

Evom Labs API documentation