FonadaLabs

Klone V2 Flash

Voice Clone · Fast

Klone V2 Flash is our faster voice-cloning model. It is a drop-in option on the same one-shot endpoint as Klone V2 Pro: add model=klone_v2_flash to the request and everything else stays the same.

Overview

Key characteristics

  • Synchronous: the audio comes back inline in the HTTP response. No job ID and no polling.
  • 12 languages: Bengali, English, Gujarati, Hindi, Kannada, Malayalam, Marathi, Odia, Punjabi, Tamil, Telugu, Urdu.
  • Up to 5,000 characters per request. The server splits and merges long text.
  • Billed at 1 credit per 8 characters, times the voice multiplier.
  • Does not use speed, duration, instruct or num_step.

Request

API endpoint

Base URL: https://api.fonada.ai
Endpoint: /v1/voice-clone/chunks
Method: POST
Content-Type: multipart/form-data (always, even with no file upload)
Authorization: Bearer YOUR_FONADA_API_KEY

curl
curl -X POST "https://api.fonada.ai/v1/voice-clone/chunks" \ -H "Authorization: Bearer YOUR_FONADA_API_KEY" \ -F "model=klone_v2_flash" \ -F "share_id=YOUR_SHARE_ID" \ -F "text=यह एक छोटा सा फ़्लैश वॉइस क्लोन टेस्ट है।" \ -F "language=hi" \ -F "output_audio_codec=mp3" \ -o voice.mp3
JavaScript (fetch)
const form = new FormData() form.append('model', 'klone_v2_flash') form.append('share_id', 'YOUR_SHARE_ID') form.append('text', 'यह एक छोटा सा फ़्लैश वॉइस क्लोन टेस्ट है।') form.append('language', 'hi') form.append('output_audio_codec', 'mp3') // optional, default wav // Do not set Content-Type yourself - the browser adds the multipart boundary. const res = await fetch('https://api.fonada.ai/v1/voice-clone/chunks', { method: 'POST', headers: { Authorization: `Bearer ${API_KEY}` }, body: form, }) if (!res.ok) { const err = await res.json() // { detail: { error, message, details? } } if (err?.detail?.error === 'rate_limit_exceeded') { const wait = err.detail.details?.retry_after_seconds ?? 5 // wait `wait` seconds, then retry } throw new Error(err?.detail?.message || 'voice clone failed') } const url = URL.createObjectURL(await res.blob()) audioElement.src = url

The body of a successful response is the raw audio. Read it as a Blob or ArrayBuffer, not JSON. Response headers: Content-Type, Content-Disposition, X-Output-Codec, X-Upstream-Time-Ms, X-Processing-Time-Ms. No credit or model headers are returned.

Parameters (multipart form)

FieldRequiredDescription
textyesText to synthesize. Non-empty, at most 5,000 characters.
languageyes (default hi)ISO code or display name. See Languages.
modelyes for FlashMust be klone_v2_flash.
output_audio_codecno (wav)See Output formats.
share_idone of 3Catalog voice ID. Works on all tiers. audio, audio_url and audio_text are ignored when it is set.
audioone of 3Reference upload (WAV/MP3/M4A/FLAC/OGG/WebM), max 10 MB and 30 s. Enterprise only.
audio_urlone of 3Public HTTPS URL of a reference clip. Enterprise only.
audio_textnoTranscript of the reference audio. Omit it to auto-transcribe. Ignored with share_id.
Send exactly one of share_id, audio or audio_url. Zero returns 400 missing_audio_source; more than one returns 400 ambiguous_audio_source. Do not send speed, duration, instruct or num_step with Flash.

Languages

Pass an ISO code or a display name (case-insensitive). Anything else returns 400 invalid_language.

CodeLanguage
bnBengali
enEnglish
guGujarati
hiHindi
knKannada
mlMalayalam
mrMarathi
orOdia
paPunjabi
taTamil
teTelugu
urUrdu

Output formats

output_audio_codecContent-TypeExtensionEncoding
wav (default)audio/wav.wavPCM WAV
mp3audio/mpeg.mp3MP3
opusaudio/ogg.opusOpus in Ogg
pcm (linear16)audio/pcm.pcmRaw headerless 16-bit PCM
mulaw (ulaw)audio/basic.ulawμ-law (telephony)
alawaudio/x-alaw-basic.alawA-law (telephony)

For web playback, mp3 or opus give the smallest payload, and wav gives the lowest latency. pcm, mulaw and alaw are headerless raw streams, so your player needs to be told the format. An unsupported value returns 400 invalid_output_audio_codec.

Billing

credits = max(round(characters / chars_per_credit × multiplier, 2), 0.01)

  • Flash: 1 credit per 8 characters by default (the per-user voice_cloning_flash_deduction_value). Pro uses 2.
  • Multiplier with share_id: the catalog voice's credit multiplier.
  • Multiplier with audio / audio_url: ×1.
  • Example: 20 characters on a 7× voice is 17.5 credits on Flash, versus 70 on Pro.
  • Credits are charged before the upstream call and refunded automatically if it fails.
  • An unknown share_id returns 404 and is not charged.

Rate limits

  • 10 requests per minute per user by default.
  • A global concurrency cap returns the same error with rate_period: "concurrent" and retry_after_seconds: 5.
  • Flash handles roughly 1.5 to 2 requests per second in total. Under heavy load, latency grows instead of erroring.
  • The wait time is only in the body. There is no Retry-After header.
429 response body
{ "detail": { "error": "rate_limit_exceeded", "message": "Rate limit exceeded. You can make 10 requests per minute.", "details": { "rate_limit": 10, "current_usage": 10, "remaining": 0, "retry_after_seconds": 37, "rate_period": "minute" } } }

Errors

All errors are JSON: {"detail": {"error": "<code>", "message": "<text>"}}

HTTPerrorCause
400empty_texttext is empty or whitespace
400text_too_longtext is longer than 5,000 characters (limit_chars and received_chars in the body); nothing is charged
400invalid_languagelanguage is not in the supported list (response includes supported_languages)
400missing_audio_sourceNo reference source supplied
400ambiguous_audio_sourceMore than one of share_id / audio / audio_url supplied
400empty_audio / audio_decode_failed / audio_too_longReference upload is empty, not valid audio, or longer than 30 s
413audio_too_largeReference upload larger than 10 MB
403enterprise_onlyaudio / audio_url on a non-enterprise tier
403invalid_api_key / inactive_api_keyBad or disabled API key
404voice_not_foundshare_id is not in the catalog
429rate_limit_exceededMore than 10 requests/minute per user, or too many requests in flight (see Rate limits)
429credits_exhaustedInsufficient credits / credit limit reached
500encoding_failedServer-side transcode to the requested codec failed
502upstream_error / upstream_unreachable / upstream_emptyUpstream cloner failed or returned no audio. Credits are refunded automatically

Python SDK

Set model="klone_v2_flash" on TTSClient.generate_audio(). Flash does not send speed or duration. As with Pro, mp3 and opus are limited to 450 characters in the SDK. See the Klone V2 Pro SDK section for installation and authentication.

Python
from fonadalabs.tts.client import TTSClient client = TTSClient(api_key="YOUR_FONADA_API_KEY") audio_bytes = client.generate_audio( text="Hello from flash cloning.", language="Hindi", model="klone_v2_flash", share_id="YOUR_SHARE_ID", output_file="flash.wav", )

FAQ

Both use POST /v1/voice-clone/chunks with the same reference sources, response headers, errors and billing formula. Flash is faster, supports 12 languages, takes the whole text in one request (the server splits and merges it), and ignores speed, duration, instruct and num_step. Only the characters-per-credit value differs: 8 for Flash, 2 for Pro.