Text-to-Speech model

Loly 3.5: Give your product a voice worth listening to

A product voice is more than an audio file. It is how reassurance sounds, how guidance keeps its rhythm and how a digital experience becomes recognisable with the screen turned off.

On this page

The contract in one screen

Every value below is read from the same modules the public routes are built on, so this table and an endpoint page can never describe two different products.

Modality
Text to speech
Transports
REST, SSE, WebSocket
Output formats
WAV, MP3, PCM16
Stream frames
PCM16 mono, 24,000 Hz
Languages
646 plus auto
Playback speed
0.5 - 1.5
Text per request
5,000 characters
Text per stream token
10,000 characters

A voice with continuity

Loly 3.5 turns text into speech while keeping the integration boundary explicit. Your application chooses the words, the voice and the transport. The model carries that intent into audio that can be stored, downloaded or played as it arrives.

Write the meaning once. Deliver it in the rhythm your experience requires.

Choose how the voice arrives

The same model serves both deliberate, file-based moments and live interfaces. Select the transport by what the listener should experience next.

Generate

REST and durable URL

Wait for a finished result and receive a URL your application can retain or distribute.

Bytes

REST and binary body

Receive the completed WAV or MP3 directly when your server wants to own the file immediately.

SSE

Progressive PCM16 events

Read base64 audio events from one POST response as synthesis progresses.

WebSocket

Binary PCM16 frames

Open a token-authorised session and play raw mono frames as they arrive.

Shape the delivery without hiding the contract

The controls are explicit request fields, not a hidden style layer. Choose a language, output container and pace, then keep advanced synthesis controls at their defaults until your listening tests give you a reason to change them.

Language

auto, code or full name

Use automatic handling or select from the published Loly 3.5 language catalog.

Format

wav or mp3

Choose a PCM container for processing or compressed audio for delivery.

Speed

0.5 to 1.5

Set the playback pace within the range enforced by the public contract.

Synthesis

cfg and diffusion steps

Advanced fields remain available when an integration needs deliberate tuning.

One credential, one voice identity

A voice key is bound to exactly one voice. That boundary makes voice selection visible in your credential design instead of letting client payloads silently switch identity. Account keys can select an allowed voice explicitly when their permissions require it.

A recognisable voice begins with permission

Custom voice enrollment is only for audio whose speaker authorised cloning. The public contract records that consent, keeps ownership scoped to the account and never returns the secret of an existing key a second time.
Upload audio for Voice Clone

Know exactly what reaches the player

Finished responses use the selected WAV or MP3 container. SSE carries base64 PCM16 little-endian mono. WebSocket binary frames carry raw PCM16 mono at the sample rate reported by the server, with JSON reserved for control events.

  • WAV or MP3 for finished output
  • Base64 PCM16 in SSE audio events
  • Raw PCM16 in WebSocket binary frames
  • Server-reported sample rate and event metadata

Let the first sentence be real

Open the Bytes endpoint for the shortest path to a playable file, or begin with WebSocket when the experience must speak while the moment is still unfolding.

Was this page helpful?

Evom Labs API documentation