The studio

Five capabilities. One continuous surface.

The Koythu studio is not a bundle of tools stitched together. Every capability shares the same voice library, the same tone controls, the same language coverage, and the same API primitives underneath.

A single canvas

Write once. Render everywhere the voice needs to go.

You do not switch between a text to speech tool and a cloning tool and a translation tool. You write a script, pick a voice, adjust tone, choose a language, and generate. Everything else - cloning, API delivery, IVR handoff - is a choice inside the same flow.

Script

Type once. Render in every voice, every tone, every language.

Aanya

Warm, calm, neutral urban

EN-IN4.2s
Rohan

Deep, formal, broadcast

HI-IN3.8s
Meera

Bright, expressive, retail

EN-IN4.0s
Your clone

Custom, consented voice

EN-IN4.5s
Capabilities

Explore the studio, surface by surface.

How the four stages nest

Write, choose, tune, generate - recursive rather than linear.

01 Write

The script is the surface

Type, paste, or feed a script through the API. Punctuation, ellipses, emphasis markers and line breaks all carry through into how the voice reads the line - a comma is a pause, a full stop is a beat, a question mark lifts.

02 Choose

Voice as a first class object

Every voice in the library carries a persona, a language coverage set, and a set of tone presets that suit it. A cloned voice slots in as a peer, not an afterthought.

03 Tune

Tone is a parameter

Rate, pitch, energy, tone family - all live sliders on the same script. You can render calm and urgent versions of the same line without re typing a word.

04 Generate

Audio, plus everything you need with it

A WAV or MP3, plus timing metadata, phoneme boundaries and language tags. Push it into an IVR, an app, a video, or the API pipeline that generated it.