Skip to main content
Those are the only valid values.
Case matters. The model interpolates whatever it is given straight into its prompt, so ananya is not the same request as Ananya — it produces audibly different audio. Anything outside the two above is rejected.

How much audio a clip is

Roughly 800 characters per minute, measured across both voices. So 10,000 characters is about 12.5 minutes.