Loly 3.5

Text-to-Speech (Bytes)

Uses the same JSON fields and permissions as batch generation, but returns the completed file directly.

POST/api/v1/tts/bytes
On this page

Code examples

curl -X POST https://studio.evomlabs.com/api/v1/tts/bytes \
  -H "Authorization: Bearer vc_sk_live_YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{"text":"Xin chào","language":"vi","format":"wav"}' \
  --output hello.wav

Authorization

Bearer credential

Send the credential in the Authorization header. Account and voice keys are accepted only where the endpoint contract permits them.

Request

FieldTypeRequiredDefaultDescription
textstringYes-Text to speak, not empty, up to 5,000 characters.
voice_idstringNo-Required with an account key. Omit with a voice key; if provided, it must match the key-bound voice.
languagestringNoautoauto, or one code or full name from the 646-language catalog.
format"mp3" | "wav"Nomp3Exactly mp3 or wav. The returned container and MIME match this value.
speednumberNo1.0Playback rate, clamped to 0.5 to 1.5. Ignored when duration is set.
durationnumber | nullNonullGenerate audio of exactly this many seconds, from 0.5 to 30. Omit it, or send null, to let the model set the length from the text. Overrides speed; out-of-range values are rejected rather than clamped.
cfg_valuenumberNo2.0How closely to follow the reference voice. A finite number from 0.0 to 4.0, including 0.
dit_stepsintegerNo10Diffusion steps. An integer from 0 to 64, including 0.

The allowance is deducted up front by text.length, including any whitespace in the string you send.

Fixed output duration

duration replaces the model's own length estimate with an exact target, which is why it overrides speed rather than combining with it. Three things to know: the saved file is not exactly that long, because postprocess_output trims silence and then pads 0.1s onto each edge — send postprocess_output: false when you need the closest match; the ceiling sits below the point where the model switches to chunked generation and re-estimates length per chunk, so a value it cannot honour is refused instead of accepted; and /sse and the WebSocket ignore the field, because they synthesise one sentence per model call and a single length has nowhere to apply. It is part of the price: billable characters are max(text length, duration × characters-per-second), so leaving duration unset costs exactly what it always did, and asking for audio longer than the text naturally needs costs more.

Response

Important

A successful response is binary audio, not JSON.
HeaderValueMeaning
Content-Typeaudio/wav | audio/mpegWAV for wav, MPEG audio for mp3.
Content-Dispositionattachment filenameSuggested oriagent.wav or oriagent.mp3 filename.
X-OriAgent-Formatwav | mp3The selected output format.
X-OriAgent-Sample-Rateinteger HzPresent only when the backend provides it.
X-OriAgent-Chars-DeductedintegerExactly text.length charged for the request.

Errors

REST failures use the documented error envelope. Handle the error code instead of matching the human-readable message.

Error codes

Was this page helpful?

Evom Labs API documentation