Klone V2 Flash
Klone V2 Flash is our faster voice-cloning model. It is a drop-in option on the same one-shot endpoint as Klone V2 Pro: add model=klone_v2_flash to the request and everything else stays the same.
Overview
Key characteristics
- Synchronous: the audio comes back inline in the HTTP response. No job ID and no polling.
- 12 languages: Bengali, English, Gujarati, Hindi, Kannada, Malayalam, Marathi, Odia, Punjabi, Tamil, Telugu, Urdu.
- Up to 5,000 characters per request. The server splits and merges long text.
- Billed at 1 credit per 8 characters, times the voice multiplier.
- Does not use
speed,duration,instructornum_step.
Request
API endpoint
Base URL: https://api.fonada.ai
Endpoint: /v1/voice-clone/chunks
Method: POST
Content-Type: multipart/form-data (always, even with no file upload)
Authorization: Bearer YOUR_FONADA_API_KEY
curl -X POST "https://api.fonada.ai/v1/voice-clone/chunks" \
-H "Authorization: Bearer YOUR_FONADA_API_KEY" \
-F "model=klone_v2_flash" \
-F "share_id=YOUR_SHARE_ID" \
-F "text=यह एक छोटा सा फ़्लैश वॉइस क्लोन टेस्ट है।" \
-F "language=hi" \
-F "output_audio_codec=mp3" \
-o voice.mp3const form = new FormData()
form.append('model', 'klone_v2_flash')
form.append('share_id', 'YOUR_SHARE_ID')
form.append('text', 'यह एक छोटा सा फ़्लैश वॉइस क्लोन टेस्ट है।')
form.append('language', 'hi')
form.append('output_audio_codec', 'mp3') // optional, default wav
// Do not set Content-Type yourself - the browser adds the multipart boundary.
const res = await fetch('https://api.fonada.ai/v1/voice-clone/chunks', {
method: 'POST',
headers: { Authorization: `Bearer ${API_KEY}` },
body: form,
})
if (!res.ok) {
const err = await res.json() // { detail: { error, message, details? } }
if (err?.detail?.error === 'rate_limit_exceeded') {
const wait = err.detail.details?.retry_after_seconds ?? 5
// wait `wait` seconds, then retry
}
throw new Error(err?.detail?.message || 'voice clone failed')
}
const url = URL.createObjectURL(await res.blob())
audioElement.src = urlThe body of a successful response is the raw audio. Read it as a Blob or ArrayBuffer, not JSON. Response headers: Content-Type, Content-Disposition, X-Output-Codec, X-Upstream-Time-Ms, X-Processing-Time-Ms. No credit or model headers are returned.
Parameters (multipart form)
| Field | Required | Description |
|---|---|---|
| text | yes | Text to synthesize. Non-empty, at most 5,000 characters. |
| language | yes (default hi) | ISO code or display name. See Languages. |
| model | yes for Flash | Must be klone_v2_flash. |
| output_audio_codec | no (wav) | See Output formats. |
| share_id | one of 3 | Catalog voice ID. Works on all tiers. audio, audio_url and audio_text are ignored when it is set. |
| audio | one of 3 | Reference upload (WAV/MP3/M4A/FLAC/OGG/WebM), max 10 MB and 30 s. Enterprise only. |
| audio_url | one of 3 | Public HTTPS URL of a reference clip. Enterprise only. |
| audio_text | no | Transcript of the reference audio. Omit it to auto-transcribe. Ignored with share_id. |
share_id, audio or audio_url. Zero returns 400 missing_audio_source; more than one returns 400 ambiguous_audio_source. Do not send speed, duration, instruct or num_step with Flash.Languages
Pass an ISO code or a display name (case-insensitive). Anything else returns 400 invalid_language.
| Code | Language |
|---|---|
| bn | Bengali |
| en | English |
| gu | Gujarati |
| hi | Hindi |
| kn | Kannada |
| ml | Malayalam |
| mr | Marathi |
| or | Odia |
| pa | Punjabi |
| ta | Tamil |
| te | Telugu |
| ur | Urdu |
Output formats
| output_audio_codec | Content-Type | Extension | Encoding |
|---|---|---|---|
| wav (default) | audio/wav | .wav | PCM WAV |
| mp3 | audio/mpeg | .mp3 | MP3 |
| opus | audio/ogg | .opus | Opus in Ogg |
| pcm (linear16) | audio/pcm | .pcm | Raw headerless 16-bit PCM |
| mulaw (ulaw) | audio/basic | .ulaw | μ-law (telephony) |
| alaw | audio/x-alaw-basic | .alaw | A-law (telephony) |
For web playback, mp3 or opus give the smallest payload, and wav gives the lowest latency. pcm, mulaw and alaw are headerless raw streams, so your player needs to be told the format. An unsupported value returns 400 invalid_output_audio_codec.
Billing
credits = max(round(characters / chars_per_credit × multiplier, 2), 0.01)
- Flash: 1 credit per 8 characters by default (the per-user
voice_cloning_flash_deduction_value). Pro uses 2. - Multiplier with
share_id: the catalog voice's credit multiplier. - Multiplier with
audio/audio_url: ×1. - Example: 20 characters on a 7× voice is 17.5 credits on Flash, versus 70 on Pro.
- Credits are charged before the upstream call and refunded automatically if it fails.
- An unknown
share_idreturns 404 and is not charged.
Rate limits
- 10 requests per minute per user by default.
- A global concurrency cap returns the same error with
rate_period: "concurrent"andretry_after_seconds: 5. - Flash handles roughly 1.5 to 2 requests per second in total. Under heavy load, latency grows instead of erroring.
- The wait time is only in the body. There is no
Retry-Afterheader.
{
"detail": {
"error": "rate_limit_exceeded",
"message": "Rate limit exceeded. You can make 10 requests per minute.",
"details": {
"rate_limit": 10,
"current_usage": 10,
"remaining": 0,
"retry_after_seconds": 37,
"rate_period": "minute"
}
}
}Errors
All errors are JSON: {"detail": {"error": "<code>", "message": "<text>"}}
| HTTP | error | Cause |
|---|---|---|
| 400 | empty_text | text is empty or whitespace |
| 400 | text_too_long | text is longer than 5,000 characters (limit_chars and received_chars in the body); nothing is charged |
| 400 | invalid_language | language is not in the supported list (response includes supported_languages) |
| 400 | missing_audio_source | No reference source supplied |
| 400 | ambiguous_audio_source | More than one of share_id / audio / audio_url supplied |
| 400 | empty_audio / audio_decode_failed / audio_too_long | Reference upload is empty, not valid audio, or longer than 30 s |
| 413 | audio_too_large | Reference upload larger than 10 MB |
| 403 | enterprise_only | audio / audio_url on a non-enterprise tier |
| 403 | invalid_api_key / inactive_api_key | Bad or disabled API key |
| 404 | voice_not_found | share_id is not in the catalog |
| 429 | rate_limit_exceeded | More than 10 requests/minute per user, or too many requests in flight (see Rate limits) |
| 429 | credits_exhausted | Insufficient credits / credit limit reached |
| 500 | encoding_failed | Server-side transcode to the requested codec failed |
| 502 | upstream_error / upstream_unreachable / upstream_empty | Upstream cloner failed or returned no audio. Credits are refunded automatically |
Python SDK
Set model="klone_v2_flash" on TTSClient.generate_audio(). Flash does not send speed or duration. As with Pro, mp3 and opus are limited to 450 characters in the SDK. See the Klone V2 Pro SDK section for installation and authentication.
from fonadalabs.tts.client import TTSClient
client = TTSClient(api_key="YOUR_FONADA_API_KEY")
audio_bytes = client.generate_audio(
text="Hello from flash cloning.",
language="Hindi",
model="klone_v2_flash",
share_id="YOUR_SHARE_ID",
output_file="flash.wav",
)FAQ
Both use POST /v1/voice-clone/chunks with the same reference sources, response headers, errors and billing formula. Flash is faster, supports 12 languages, takes the whole text in one request (the server splits and merges it), and ignores speed, duration, instruct and num_step. Only the characters-per-credit value differs: 8 for Flash, 2 for Pro.
