Generate
REST and durable URL
Wait for a finished result and receive a URL your application can retain or distribute.
Text-to-Speech model
A product voice is more than an audio file. It is how reassurance sounds, how guidance keeps its rhythm and how a digital experience becomes recognisable with the screen turned off.
Every value below is read from the same modules the public routes are built on, so this table and an endpoint page can never describe two different products.
Loly 3.5 turns text into speech while keeping the integration boundary explicit. Your application chooses the words, the voice and the transport. The model carries that intent into audio that can be stored, downloaded or played as it arrives.
Write the meaning once. Deliver it in the rhythm your experience requires.
The same model serves both deliberate, file-based moments and live interfaces. Select the transport by what the listener should experience next.
REST and durable URL
Wait for a finished result and receive a URL your application can retain or distribute.
REST and binary body
Receive the completed WAV or MP3 directly when your server wants to own the file immediately.
Progressive PCM16 events
Read base64 audio events from one POST response as synthesis progresses.
Binary PCM16 frames
Open a token-authorised session and play raw mono frames as they arrive.
The controls are explicit request fields, not a hidden style layer. Choose a language, output container and pace, then keep advanced synthesis controls at their defaults until your listening tests give you a reason to change them.
auto, code or full nameUse automatic handling or select from the published Loly 3.5 language catalog.
wav or mp3Choose a PCM container for processing or compressed audio for delivery.
0.5 to 1.5Set the playback pace within the range enforced by the public contract.
cfg and diffusion stepsAdvanced fields remain available when an integration needs deliberate tuning.
A voice key is bound to exactly one voice. That boundary makes voice selection visible in your credential design instead of letting client payloads silently switch identity. Account keys can select an allowed voice explicitly when their permissions require it.
A recognisable voice begins with permission
Finished responses use the selected WAV or MP3 container. SSE carries base64 PCM16 little-endian mono. WebSocket binary frames carry raw PCM16 mono at the sample rate reported by the server, with JSON reserved for control events.
Open the Bytes endpoint for the shortest path to a playable file, or begin with WebSocket when the experience must speak while the moment is still unfolding.
Was this page helpful?