Skip to content

blab-pro

What is this model?

Premium text-to-speech built on VoxCPM2, served on our own GPUs via nano-vllm. Produces 48 kHz studio-grade, context-aware speech across 30 languages (including Turkish), with instant voice cloning from a reference clip, prompt-based voice design, and natural-language style direction.

Capability & type

CapabilityText to Speech (TTS)
Model typeText-to-speech (TTS) (type=7)
Base modelVoxCPM2 (nano-vllm)
Providerarpanet (local / self-hosted)
Streaming✓ supported

Endpoint

POST /v1/audio/speech

Base URL: https://app.qevron.ai/v1

Authentication

Every request needs the Authorization: Bearer <KEY> header.

Authorization: Bearer <KEY>

Request schema & parameters

ParameterTypeRequiredDescription
modelstringblab-pro
inputstringText to synthesize
voicestringVoice id
response_formatstringwav

Response schema

audio/wav (binary) — the response body is the audio stream.

Streaming

Add "stream": true; the server replies incrementally over Server-Sent Events.

Code examples

curl

bash
curl https://app.qevron.ai/v1/audio/speech \
  -H "Authorization: Bearer $QEVRON_KEY" -H "Content-Type: application/json" \
  -d '{"model": "blab-pro", "input": "Merhaba dünya", "voice": "default", "response_format": "wav"}' \
  --output speech.wav

Python

python
from openai import OpenAI
client = OpenAI(api_key="$QEVRON_KEY", base_url="https://app.qevron.ai/v1")
r = client.audio.speech.create(model="blab-pro", input="Merhaba dünya", voice="default")
r.stream_to_file("speech.wav")

Pricing

InputOutputUnit
$0.00121K tokens

Rate limits

RPM/TPM are governed by your group quota. Check your current quota with: GET /v1/dashboard/billing/quota.

Notes & quirks

  • Style direction: pass a short natural-language instruction (e.g. "cheerful, slightly faster"). There is no numeric emotion knob.
  • Voice design: voice_description generates a brand-new voice with no reference audio; seed makes a candidate reproducible.
  • Instant cloning: POST /v1/audio/voices?model=blab-pro (multipart label+file, optional transcript) — ready in seconds. An exact transcript upgrades the clone to prompt-continuation ("ultimate") mode for sharper similarity.
  • Streaming: stream: true yields SSE speech.audio.delta; response_format: "pcm" is a 24 kHz int16 stream (realtime agents).
  • Identical requests reproduce identical audio (a deterministic request-derived seed); change seed for a fresh take.

Limits & constraints

  • File outputs are wav/mp3 (48 kHz); opus/aac/flac are not supported.
  • Clone references must be 3–60 s; heavily processed/jingled recordings may be rejected by the quality probe.
  • The blab-expressive and blab-tts ids are permanent aliases of this model.

Qevron — AI gateway. Arpanet / OpenAI / Anthropic / Gemini compatible.